arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型表示的统一视角:从填充符-角色结构到机制可解释性

A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

Zhang Enyan, R. Thomas McCoy

arXiv 2608.29034首次发表:更新:

发表机构

Yale University; Wu Tsai Institute(耶鲁大学; 吴蔡研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以张量积表示(TPRs)为统一假设,证明其可统一加法类比等多种语言模型可解释性方法,构建的变体与标准变体性能相当,为神经网络本质的统一阐释奠定基础。

AI 中文摘要

针对语言模型的解释已提出大量方法,为其内部运行机制提供了重要洞见,但不同方法及所得洞见相对孤立:语言模型的底层结构究竟是什么,才能催生所有这些解释?本研究提出用张量积表示(Tensor Product Representations, TPRs)作为统一假设。TPRs为向量空间中组合结构的表示提供了具体方案——即填充符-角色绑定。我们从数学和实证两方面证明,TPRs可统一多种现有可解释性方法:加法类比、线性探测、稀疏自编码器及激活修补。数学上,这些方法均可从TPRs推导得出;实证上,我们将推导结果应用于从小型玩具模型到大语言模型(LLMs)的多种不同模型,构建上述每种可解释性方法的实例,这些构建出的变体与标准变体性能相当。我们认为本研究是迈向可解释性理想状态的一步:即对神经网络本质的统一阐释,不仅得到个别观察结果的佐证,也能解释各观察结果间的关联。

英文摘要

A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language models be, such that they give rise to all our interpretations? In this work, we propose using Tensor Product Representations (TPRs) as a unifying hypothesis. TPRs give a concrete proposal for how compositional structure could be represented in vector space --- as filler-role bindings. We show, both mathematically and empirically, that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching. Mathematically, we show that these methods can all be derived from TPRs. Empirically, we apply the derivations to a range of different models --- from small toy models to LLMs --- to construct instances of each of the above interpretability methods; these constructed variants perform comparably to their standard variants. We view this work as a step toward what interpretability will ideally provide: a unified account of the nature of neural networks, corroborated not just by individual observations but also by an explanation of the connections between them.

Comments32 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑