arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型激活占据特权纠错盆地

Language Model Activations Inhabit Privileged Error-Correcting Basins

Matthew Finlayson, Francisco Pernice, Eric Todd, Amir Zur, Daniel Wurgaft, Fenil R. Doshi, Vasudev Shyam, Matt Feiszli, Satchel Grant, Lucius Bushnaq, Tal Haklay, Usha Bhalla, Matthew Kowal, Thomas Fel, Jack Merullo, Atticus Geiger, Xiang Ren, Owen Lewis, Ekdeep Singh Lubana

arXiv 2610.04183首次发表:更新:

发表机构

Goodfire; University of Southern California; Massachusetts Institute of Technology; Northeastern University; Stanford University; Harvard University(Goodfire; 南加州大学; 麻省理工学院; 东北大学; 斯坦福大学; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过探测语言模型激活空间的几何结构,发现激活位于特权纠错盆地中,并利用自适应引导强度跨盆地传输激活,显著提升跨语言引导效果。

AI 中文摘要

语言模型表现出显著的鲁棒性,即使其激活受到线性引导等干预的扰动,仍能继续生成连贯的文本。我们假设这种鲁棒性是被动动力学的结果,即前向传播中的约束机制将激活引向产生连贯输出的“良好”区域。为了研究这些假设的纠错机制,我们通过观察模型层对低维曲线的作用来探测语言模型激活空间的几何结构。在此过程中,我们发现模型激活发生在一组不同的吸引盆地内,这些盆地在几何上将自然激活与分布上相似的合成激活区分开来。将这一视角应用于语言模型引导,我们观察到沿语义引导方向存在特征特定的盆地,并发现引导会在这些盆地之间移动激活。为了展示这种几何结构在神经计算中的积极作用,我们表明自适应地调节引导强度以在盆地间传输激活可改善跨语言引导,与固定强度引导相比,显著提高了从目标语言采样标记的概率。我们的发现确立了激活空间几何分析作为解释和控制语言模型的一种有前景的方法。

英文摘要

Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward "good" regions that produce coherent outputs. To investigate these hypothesized error-correcting mechanisms, we probe the geometry of language model activation space by observing the action of model layers on low-dimensional curves. In doing so, we discover that model activations occur within a cluster of distinct attracting basins, which differentiate natural activations geometrically from distributionally similar synthetic activations. Applying this lens to language model steering, we observe feature-specific basins along semantic steering directions, and find that steering moves activations between these basins. To demonstrate the active role of this geometry in neural computation, we show that adaptively modulating steering strength to transport activations across basins improves inter-language steering, significantly increasing the probability of sampling tokens from the target language compared to fixed-strength steering. Our findings establish analysis of activation space geometry as a promising approach to interpreting and controlling language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑