发表机构
Yale University; Wu Tsai Institute(耶鲁大学; 吴蔡研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以张量积表示(TPRs)为统一假设,证明其可统一加法类比等多种语言模型可解释性方法,构建的变体与标准变体性能相当,为神经网络本质的统一阐释奠定基础。
AI 中文摘要
针对语言模型的解释已提出大量方法,为其内部运行机制提供了重要洞见,但不同方法及所得洞见相对孤立:语言模型的底层结构究竟是什么,才能催生所有这些解释?本研究提出用张量积表示(Tensor Product Representations, TPRs)作为统一假设。TPRs为向量空间中组合结构的表示提供了具体方案——即填充符-角色绑定。我们从数学和实证两方面证明,TPRs可统一多种现有可解释性方法:加法类比、线性探测、稀疏自编码器及激活修补。数学上,这些方法均可从TPRs推导得出;实证上,我们将推导结果应用于从小型玩具模型到大语言模型(LLMs)的多种不同模型,构建上述每种可解释性方法的实例,这些构建出的变体与标准变体性能相当。我们认为本研究是迈向可解释性理想状态的一步:即对神经网络本质的统一阐释,不仅得到个别观察结果的佐证,也能解释各观察结果间的关联。
英文摘要
A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language models be, such that they give rise to all our interpretations? In this work, we propose using Tensor Product Representations (TPRs) as a unifying hypothesis. TPRs give a concrete proposal for how compositional structure could be represented in vector space --- as filler-role bindings. We show, both mathematically and empirically, that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching. Mathematically, we show that these methods can all be derived from TPRs. Empirically, we apply the derivations to a range of different models --- from small toy models to LLMs --- to construct instances of each of the above interpretability methods; these constructed variants perform comparably to their standard variants. We view this work as a step toward what interpretability will ideally provide: a unified account of the nature of neural networks, corroborated not just by individual observations but also by an explanation of the connections between them.
Comments32 pages, 6 figures