发表机构
Institute for Computational & Mathematical Engineering, Stanford University; Department of Computer Science, Stanford University; Department of Computer Science and Engineering, University at Buffalo(斯坦福大学计算与数学工程研究所; 斯坦福大学计算机科学系; 布法罗大学计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型中MLP层以最优速率存储事实知识的现象,开发首个与Transformer兼容的事实存储MLP闭式构造,具有最优存储缩放等特性,所需参数更少,还能用于事实召回任务及实现模块化事实编辑。
AI 中文摘要
大语言模型(LLMs)在其参数中存储事实知识。虽然近期工作表明这种知识存在于多层感知器(MLP)层中,但现有的LLMs中事实存储的建设性和机制性可解释性模型未能解释它们以信息理论最优速率存储事实这一惊人的经验现象。在这项工作中,我们对这一现象进行了理论阐述。我们开发了首个与Transformer兼容的事实存储MLP闭式构造,它满足在LLMs中凭经验观察到的三个属性:(i)实现最优事实存储缩放,(ii)处理任意输入/输出几何结构,(iii)在Transformer内部工作。我们工作的关键是分析MLP的解码余量,而先前工作仅研究MLP事实存储。在各向同性嵌入下,我们的构造实现了信息理论最优存储容量缩放,并且在匹配事实数量时所需参数比先前构造少10 - 104倍。对于任意键和值嵌入,我们表明我们的构造达到相同的存储容量缩放,至多取决于嵌入几何结构的惩罚因子。此外,我们证明我们构造的MLP可用于Transformer块内的事实召回任务,在匹配事实数量时所需参数比先前构造少15 - 63倍。最后,作为概念验证,我们表明事实存储MLP通过用新的MLP替换Transformer的MLP实现模块化事实编辑。
英文摘要
Large language models (LLMs) store factual knowledge in their parameters. While recent work has shown that this knowledge resides in MLP layers, existing constructive and mechanistic interpretability models of fact-storage in LLMs fail to explain the surprising empirical phenomenon that they store facts at an information-theoretically optimal rate. In this work, we develop a theoretical account of this phenomenon. We develop the first Transformer-compatible fact-storing MLP closed-form construction that satisfies the following three properties empirically observed in LLMs: it (i) attains optimal fact storage scaling, (ii) handles arbitrary input/output geometries, and (iii) works inside Transformers. Key to our work is to analyze the decoding margin of MLPs, whereas prior work only studies MLP fact storage. Under isotropic embeddings, our construction achieves information-theoretically optimal storage capacity scaling and requires $10$-$104\times$ fewer parameters at matched fact count than prior constructions. For arbitrary key and value embeddings, we show that our construction attains the same storage capacity scaling, up to penalization factors depending on the embedding geometries. Moreover, we demonstrate that our constructed MLPs can be used within Transformer blocks for factual recall tasks at optimal capacity scaling, requiring $15$-$63\times$ fewer parameters at matched fact count than prior constructions. Finally, as a proof-of-concept, we show that fact-storing MLPs enable modular fact editing by swapping a Transformer's MLP with a new one.