arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度学习模型也会回忆特征

Deep Learning Models Also Recall Features

Pierre Beckmann

arXiv 2608.20970首次发表:更新:

发表机构

École Polytechnique Fédérale de Lausanne (EPFL); Idiap Research Institute(洛桑联邦理工学院; Idiap研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出深度学习模型存在特征回忆这一通用操作,定义了该概念、证明其适用于多种架构并与特征组合对比,为理解深度学习提供了新工具,也为机械可解释性研究指明了实证方向。

AI 中文摘要

近期机械可解释性领域的研究已探究了大型语言模型如何回忆存储在其权重中的事实。本文提出,事实回忆指向更广泛的内容:深度学习模型中的一类通用操作,我将其命名为特征回忆。核心观察是,线性投影可被解读为按输入激活缩放后检索存储信息。我定义了特征回忆,证明其适用于多种架构,并将其与已确立的特征组合范式进行对比。我还探讨了如何从机制上识别特征回忆的案例。该论述为哲学家提供了理解深度学习的新概念工具,并为机械可解释性研究指明了实证方向。

英文摘要

Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which I call feature recall. The core observation is that a linear projection can be read as retrieving stored information scaled by input activations. I define feature recall, show it applies across architectures, and contrast it with the established paradigm of feature combination. I also consider how cases of feature recall might be mechanistically identified. The account gives philosophers a new conceptual tool for understanding deep learning, and points to empirical directions for mechanistic interpretability research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑