Transformer残差流中的几何与行为分层
Geometric and Behavioral Stratification in Transformer Residual Streams
浏览论文内容
中文总结 AI 辅助
该研究发现Transformer残差流的预测方向是特权锚点,其附近区域高度结构化,远端区域平缓,破坏近端方差方向会导致任务框架转移,为高维计算与线性读出的共存提供了几何解释。
中文摘要 AI 辅助
经过训练的Transformer模型会形成特权基:其坐标轴的统计特性与残差流其余部分不同。但这类基会选择何种方向?我们研究了预测方向,即模型当前要预测的token的解嵌入方向,发现其可作为内容定义的特权锚点。相对于该锚点测量,残差流变化在几何和行为上按与预测的接近程度分层。这种分层在所有测试的18个模型(包括密集模型和混合专家模型,参数规模7B至120B,基础模型和指令微调模型)中均成立。狭窄的、尺度不变的预测接口集中了与读出相关的结构,而占预测远端的互补部分则随模型规模扩大而扩展。由于预测方向几乎与主方差轴正交,基于方差的分析只能部分恢复该组织,且随着提示异质性的增加,不足程度会增大。锚定揭示了陡峭的几何梯度:预测近端区域高度结构化,会对相关提示进行聚类,而互补部分则更平缓,会对提示组进行反区分。该接口是狭窄的切片,但在功能上具有决定性:破坏最接近预测的方差方向会导致立即发散和频繁的任务框架转移;破坏下一层则会延迟发散并保留框架。互补部分在方向上与读出的对齐程度较弱,但在因果和时间上具有承载作用,且行为由方向而非幅度驱动。这些结果确立了预测方向是一种特权锚点,与之前描述的坐标轴不同,并为高维计算如何与线性读出共存提供了几何解释。
英文摘要
Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.