arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果残差变压器中中间丢失现象的伴随灵敏度框架

An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers

Cheng Huan, Hongwei Yuan

arXiv 2607.17696首次发表:更新:

AI 中文总结

该研究为因果残差变压器的位置影响构建伴随灵敏度框架,通过定义相关密度推导其演化,给出定理及分解,探讨多种机制对U形轮廓的作用,提出边界优势及相关诊断或正则化器,模拟表明各干预可控,相关平衡或加权不必然降低诊断。

AI 中文摘要

我们为因果残差变压器中的位置影响开发了一个伴随灵敏度框架,并将无条件分析结果与有条件边界形状结论分开。主要的无条件定理是层控制在\(L^1\)中收敛的残差到深度流估计,辅以有限令牌到沃尔泰拉注意力估计,该估计明确控制因果端点附近的第一个单元。我们定义了归一化伴随能量影响密度,并推导了其沿全批量梯度流的精确演化。伴随允许精确的生成项分解为残差传输、非局部沃尔泰拉和局部通道,包括所有协方差交叉项。因果掩码可以放大早期位置灵敏度,残差恒等路径可以传输右局部化终端偏差,但单独的任何一种机制都不会强制形成U形轮廓。因此,我们在可独立检查的能量、相关性和局部通道边界下陈述边界优势;这些条件是充分而非必要的。有限令牌影响平衡、位置重新加权和任务对齐可观测性作为诊断或正则化器呈现,具有明确的微分要求、计算成本和局限性。受控模拟表明,每种干预都控制其指定的代理,而可观测性平衡或外环重新加权不一定会单调降低基于影响的中间丢失诊断。

英文摘要

We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions. The principal unconditional theorem is the residual-to-depth-flow estimate for layer controls converging in $L^1$, complemented by a finite-token-to-Volterra attention estimate that explicitly controls the first cells near the causal endpoint. We define a normalized adjoint-energy influence density and derive its exact evolution along full-batch gradient flow. The adjoint admits an exact generator-term decomposition into residual transmission, nonlocal Volterra, and local channels, including all covariance cross terms. Causal masking can amplify early-position sensitivity and residual identity paths can transmit a right-localized terminal bias, but neither mechanism alone forces a U-shaped profile. We therefore state boundary advantages under independently checkable energy, correlation, and local-channel bounds; these conditions are sufficient rather than necessary. Finite-token influence balancing, positional reweighting, and task-aligned observability are presented as diagnostics or regularizers with explicit differentiation requirements, computational costs, and limitations. Controlled simulations illustrate that each intervention controls its designated surrogate, while observability balance or outer-loop reweighting need not monotonically reduce the influence-based Lost-in-the-Middle diagnostic.

Comments41 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑