真相方向剖析:小语言模型中依赖知识的维度、关系定律与收敛类别几何
The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Shared Category Geometry in Small Language Models
浏览论文内容
中文总结 AI 辅助
研究小语言模型中真相方向,通过无训练定向探针及多模型实验,探讨真相维度与知识的关系、架构组件作用及方向混合情况,揭示关系定律与知识门控定律,表明混合几何属知识领域。
中文摘要 AI 辅助
Bürger等人(2024年)证明,大语言模型中的真相表示在语句极性上具有普遍性,但存在于多维子空间中。本文沿着三个问题扩展了该框架:子空间的维度如何依赖于模型的知识,哪个架构组件构建了真相方向,以及该方向是何种混合。第一部分通过无训练定向探针表明真相维度依赖知识,信号在行为已知事实时集中于单轴,知识减少或材料异质性增加时扩散,七个证伪实验支持一维解读,相同分解可恢复监督极性方向。第二部分出现关系定律,注意力传播未写入的真相框架,前馈网络反对当前块的框架,真实的峰值后衰减因果归因于所有四个模型的SwiGLU值流;类别真相轴形成语义有符号排列,跨家族收敛(Mantel p = 0.0009)。在三个尺度上的压力测试暴露了类别方向的符号不稳定性,用声明的光谱共识规范修复并锐化收敛为知识门控定律,从密集混合知识材料上的+0.42增强到知识受限类别上的+0.74:混合几何属于知识领域而非架构。
英文摘要
Bürger et al.\ (2024) demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace. The truth value of a statement is linearly readable from a residual stream of language model, but it is not clear how much of that representation fits on a single direction, which component builds it, or what it is made of. We conducted a study based on these questions, with one instrument: a training-free axis, the dominant direction of the singular value decomposition (SVD) of hidden-state differences over true/false minimal pairs, identified without labels up to one global sign. Extensive evaluation across 14 models from 6 diverse architectural families (including MoE), read and extract at cost $O(d)$ per token. We close with a pre-registered prediction on whether the arrangement extends to categories whose truth is computed rather than retrieved.