arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37717cs.LGcs.CL

Transformer中隐藏轨迹的预测几何

Predictive Geometry of Hidden Trajectories in Transformers

  • University of Luxembourg(卢森堡大学)
  • London Institute for Mathematical Sciences(伦敦数学科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State

中文总结 AI 辅助

本文通过逐层损失几何和拉回Fisher算子揭示Transformer隐藏轨迹的预测结构,提出逐词曲率分数,支持剪枝、秩分配和蒸馏,提升模型压缩效果。

中文摘要 AI 辅助

仅解码器Transformer仅通过终端的下一词预测损失进行训练,然而该损失通过固定的下游计算约束了每一个中间隐藏状态。我们通过研究逐层损失到目标函数(即通过剩余Transformer模块继续候选隐藏状态所获得的终端损失)来形式化这一约束。在成功的验证轨迹附近,我们证明这些函数的局部二阶几何由隐藏状态空间上的拉回Fisher算子控制,直至低损失残差项。其谱识别出输出敏感方向和近似预测零方向,从而得到残差流的局部可观测子空间。对于因果Transformer,相同的几何结构引出一个逐词曲率分数:即目标logits对每个词隐藏状态扰动的Fisher加权敏感性。该分数在目标的因果祖先集之外消失,并受下游Jacobian耦合控制,使其成为注意力幅度的损失感知替代方案。我们使用无矩阵的Jacobian-向量和向量-Jacobian乘积来估计这些量,并在WikiText、OpenWebText和FineWeb上的仅解码器语言模型中评估它们。实验上,所诱导的几何结构预测扰动敏感性,支持非均匀的逐层秩分配,产生有竞争力的结构化词元剪枝信号,并在添加到更强的自回归蒸馏目标(如反向KL和偏斜KL)时改善低秩学生恢复。这些结果支持对Transformer计算的预测几何观点:在成功轨迹附近,终端损失诱导出一组薄且各向异性的输出相关隐藏状态方向,这些方向可以被测量并用于压缩和蒸馏。

英文摘要

Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local second-order geometry of these functions is governed, up to low-loss residual terms, by a pullback Fisher operator on hidden-state space. Its spectrum identifies output-sensitive directions and approximately prediction-null directions, yielding a local observable subspace of the residual stream. For causal transformers, the same geometry induces a tokenwise curvature score: a Fisher-weighted sensitivity of the target logits to perturbations of each token's hidden state. This score vanishes outside the causal ancestor set of the target and is controlled by downstream Jacobian couplings, making it a loss-aware alternative to attention magnitude. We estimate these quantities using matrix-free Jacobian-vector and vector-Jacobian products and evaluate them across decoder-only language models on WikiText, OpenWebText, and FineWeb. Empirically, the induced geometry predicts perturbation sensitivity, supports nonuniform layerwise rank allocation, yields competitive structured token-pruning signals, and improves low-rank student recovery when added to stronger autoregressive distillation objectives such as reverse KL and skew KL. These results support a predictive-geometric view of transformer computation: near successful trajectories, the terminal loss induces a thin, anisotropic set of output-relevant hidden-state directions that can be measured and exploited for compression and distillation.

↑