arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

离轴,出于目的:Transformer在何处计算概念以及为何如此计算

Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

Mark Oskin

arXiv 2608.10251首次发表:更新:

发表机构

University of Washington; School of Computer Science and Engineering(华盛顿大学; 计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示Transformer分两阶段计算,概念阶段采用与读取轴近乎正交的子空间,强制该几何结构可提升稀疏旋转的收敛性,且不损害语言建模基准性能。

AI 中文摘要

Transformer的答案存在于一个轴上:其非嵌入层(unembedding)读取的方向。其中间状态大多不在该轴上,而这种离轴位置通常被视为解释的障碍。我们证明它具有功能性。一个12层模型分两个阶段计算:第一阶段,每个子层写入一个与读取方向近乎正交的子空间,各层的注意力与该轴的夹角为75至96度;将注意力的值移到读取轴上造成的损害比匹配的随机旋转大64至84倍,且损害完全出现在跨词元混合中:该子空间将组合与词汇隔离开,其下的框架随深度刚性变化。第二阶段,答案在后期沿轴出现,通过加法而非将累积内容转向读取轴产生。若像早期退出训练那样将每层压到读取轴上,在困惑度、LAMBADA和BLiMP基准上与基线匹配,同时将概念阶段工作空间从约25个有效维度降至14个,而这些基准均未记录该变化。该几何结构也可被强制施加,但并非通过主动要求:通过损失函数指定它是一场概率游戏,8个随机种子中有6个崩溃,因为被要求置零读取投影的模型最廉价的服从方式是丢弃维度;在阶段边界插入一个固定旋转则能成功,达到基线质量。周围权重可吸收的稀疏旋转在所有9个种子上收敛,而普通训练仅5个种子收敛;旋转类型无关紧要,13种不同旋转的25次运行均达到相同质量,来自不同种子的两个基线在读取轴一致的同时,在近乎正交的框架中保留其概念,这种自由度是可用的:训练前指定的随机基在概念阶段被采用,质量不变。

英文摘要

A transformer's answer lives on one axis: the direction its unembedding reads. Its intermediate states largely do not, and that off-axis position is usually treated as an obstacle to interpretation. We show it is functional. A 12-layer model computes in two phases. Through the first, every sublayer writes into a subspace held near-orthogonal to the read-out, attention 75 to 96 degrees off it at every depth. Moving attention's values onto the read-out is 64 to 84 times more damaging than a matched random rotation, and the damage is entirely in cross-token mixing: the subspace insulates composition from the vocabulary. Beneath it the frame itself turns rigidly with depth. In the second phase the answer arrives on-axis, late, and by addition rather than by turning accumulated content onto the read-out. Pressing every layer onto the read-out instead, as training for early exit does, matches the baseline on perplexity, LAMBADA and BLiMP while cutting the concept-phase workspace from about twenty-five effective dimensions to fourteen, a change none of those benchmarks register. The geometry can also be imposed, though not by asking for it. Prescribing it through the loss is a lottery: six of eight seeds collapse, because a model told to null its read-out projection obeys most cheaply by discarding dimensions. Inserting one fixed rotation at the phase boundary lands it instead, at baseline quality. A sparse rotation the surrounding weights can absorb converges on all nine seeds, against five of nine for ordinary training. Which rotation is immaterial: twenty-five runs across thirteen distinct ones reach the same quality, and two baselines from different seeds hold their concepts in near-orthogonal frames while agreeing on their read-outs. That freedom is usable: a basis drawn at random and prescribed before training is adopted across the concept phase, with quality unchanged.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑