arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04661cs.LGcs.FLstat.ML

图灵机的可解释性

Interpretability for Turing Machines

  • University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

Billy Snikkers, Rumi Salazar, Daniel Murfet, Will Troiani

AI总结:

该研究将神经网络的可解释性技术敏感度应用于图灵机,通过理论证明与DFAs实证研究,证实可通过该技术及相关方法恢复图灵机的算法特征。

AI中文摘要:

我们证明,针对神经网络开发的可解释性技术——敏感度(susceptibilities),可通过探测由Murfet和Troiani(arXiv:2504.08075)引入的带噪图灵机学习问题的局部损失景观,识别图灵机中算法结构的存在。我们证明,图灵机所实现算法中的对称性与路径分离,会在其敏感度矩阵中诱导出置换对称性与低秩块。我们在一组确定有限自动机(DFAs)上进行了实证研究,证明可通过主成分分析与聚类方法在敏感度空间中恢复算法特征。

英文摘要:

We show that susceptibilities, an interpretability technique developed for neural networks, can identify the presence of algorithmic structure in Turing machines by probing the local loss landscape of a learning problem for noisy Turing machines introduced by Murfet and Troiani (arXiv:2504.08075). We prove that symmetries and path separation in the algorithm implemented by a Turing machine induce permutation symmetries and low-rank blocks in its susceptibility matrix. We study this empirically on a set of deterministic finite automata (DFAs) and demonstrate that algorithmic features can be recovered by principal component analysis and clustering methods in susceptibility space.

补充信息

↑