图灵机的可解释性
Interpretability for Turing Machines
- University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究将神经网络的可解释性技术敏感度应用于图灵机,通过理论证明与DFAs实证研究,证实可通过该技术及相关方法恢复图灵机的算法特征。
AI中文摘要:
我们证明,针对神经网络开发的可解释性技术——敏感度(susceptibilities),可通过探测由Murfet和Troiani(arXiv:2504.08075)引入的带噪图灵机学习问题的局部损失景观,识别图灵机中算法结构的存在。我们证明,图灵机所实现算法中的对称性与路径分离,会在其敏感度矩阵中诱导出置换对称性与低秩块。我们在一组确定有限自动机(DFAs)上进行了实证研究,证明可通过主成分分析与聚类方法在敏感度空间中恢复算法特征。
英文摘要:
We show that susceptibilities, an interpretability technique developed for neural networks, can identify the presence of algorithmic structure in Turing machines by probing the local loss landscape of a learning problem for noisy Turing machines introduced by Murfet and Troiani (arXiv:2504.08075). We prove that symmetries and path separation in the algorithm implemented by a Turing machine induce permutation symmetries and low-rank blocks in its susceptibility matrix. We study this empirically on a set of deterministic finite automata (DFAs) and demonstrate that algorithmic features can be recovered by principal component analysis and clustering methods in susceptibility space.