发表机构
Shadan Women’s College of Engineering and Technology(沙丹女子工程技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探究新注意力机制在百万 token 上下文中能否修复注意力汇聚问题,通过构建 SinkProbe 套件进行测量,发现训练目标导致汇聚,门控机制效果未复现,且各指标独立变化。
AI 中文摘要
长上下文语言模型现在宣称支持百万 token 的窗口,但有两个习惯限制了窗口的使用范围。注意力头在没有有用信息可读时,仍会将预算花费在第一个 token 上,这被称为注意力汇聚(attention sink),而且事实在上下文中的位置会影响模型能否找到它。门控注意力在 NeurIPS 2025 上将首个 token 的注意力从 46.7% 降至 4.8%,Kimi K3 将该想法与 Kimi Delta Attention 和 Attention Residuals 结合,支持百万 token 窗口,比这些诊断报告的范围高出八倍。本文探讨这一修复能否在如此大的跨度下依然有效。我们构建了 SinkProbe 套件,用于测量汇聚质量、大规模激活、位置解析召回率和近期性差距,并将其应用于四个仅在 token 混合和深度上有所差异的小型模型。我们得出三个结果:训练目标而非架构导致了汇聚现象;门控机制在我们的规模下未能复现其已发表的效果;汇聚质量、激活和位置偏差独立变化。代码、数据和测量协议已在此 https URL 发布。
英文摘要
Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used. Attention heads with nothing useful to read still spend their budget on the first token, which is called the attention sink, and where a fact sits in the context changes whether the model finds it. Gated attention cut first token attention from 46.7 percent to 4.8 percent at NeurIPS 2025, and Kimi K3 pairs that idea with Kimi Delta Attention and Attention Residuals behind a one million token window, eight times past the range where these diagnostics have been reported. This paper asks whether the fix survives that jump. We build SinkProbe, a suite that measures sink mass, massive activation, position resolved recall and the recency gap, and apply it to four small models that differ only in how they mix tokens and depth. Three results follow. The training objective produces the sink, not the architecture. Gating did not reproduce its published effect at our scale. Sink mass, activations and position bias moved independently. Code, data and the measurement protocol are released at https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window
CommentsExperimental study of attention sinks, long-context recall, and million-token context behavior. Code and measurement protocol are available at https://github.com/sararizwan7/Attention-Mechanisms-in-1M-Context-Window