arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

残差流的有效深度

The Residual Stream's Effective Depth

Barak Gahtan, Ido Galil, Alex M. Bronstein

arXiv 2609.31098首次发表:更新:

发表机构

Technion Israel Institute of Technology; Nvidia; ISTA Institute of Science and Technology Austria(以色列理工学院; 英伟达; 奥地利科学技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出有效深度($\Deff$)诊断指标,量化Transformer残差流中表示相似性随层距离的衰减,发现多数模型低于理论参考值,表明残差更新相关而非深度未用,并验证其作为全局累积状态诊断的有效性。

AI 中文摘要

我们引入了“有效深度”($\Deff$),这是一种标量诊断指标,将Transformer的逐层残差流视为离散时间过程,度量表示相似性随层距离的衰减方式,并将该轮廓聚合为一个数字。在十六个仅解码器语言模型中,$\Deff$区分了残差累积的结构性后果与经验性后果:即使最大多样性的正交更新也有闭式参考值$F_L = 2L/(L+1)<2$,然而十六个默认测量中有十五个低于$F_L$(Qwen3.5:32–44%,OLMo-2:40–41%,Pythia:23–28%)。匹配的参考值表明,这一差距并非由持续的初始状态或更新大小不平衡造成,而主要是相关残差更新的校准特征,而非深度未被使用的证据。对称位置0、令牌归一化和主成分(top-PC)控制表明,该状态不能归结为BOS或top-PC伪影:唯一高于参考值的默认异常值也进入同一状态,并且在令牌归一化或去除top-1主成分后,所有十六个模型均低于参考值。中间检查点显示,该状态在OLMo-2中早期建立,并在5T令牌期间保持稳定,而Pythia-1.4B则遵循不同的递减轨迹。受控的残差携带干预支持该机制,并且$\Deff$最好被解读为一种“全局”累积状态诊断指标,而非能力评分或剪枝方法。

英文摘要

We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference $F_L = 2L/(L+1)<2$, yet fifteen of sixteen default measurements lie below $F_L$ (Qwen3.5: 32--44\%, OLMo-2: 40--41\%, Pythia: 23--28\%). Matched references show that the gap is not caused by the persistent initial state or update-size imbalance, but is largely a calibrated signature of correlated residual updates rather than evidence that depth is unused. Symmetric position-0, token-normalisation, and top-PC controls show the regime is not reducible to BOS or top-PC artefacts: the lone above-reference default outlier joins the same regime, and all sixteen models are sub-reference after token-normalisation or top-1-PC removal. Intermediate checkpoints show that the regime is established early in OLMo-2 and stable through 5T tokens, while Pythia-1.4B follows a distinct decreasing trajectory. A controlled residual-carry intervention supports the mechanism, and $\Deff$ is best read as a \emph{global} accumulated-state diagnostic, not as a capability score or pruning method.

CommentsAccepted at the 17th Asian Conference on Machine Learning (ACML 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑