发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出NAMOH稀疏注意力机制,通过每词元激活K个头的设计,实现参数缩放直接驱动上下文缩放,在相同参数下超越全激活模型并提升长上下文推理效率。
AI 中文摘要
缩放注意力参数可以提高语言模型的质量,但在长上下文中保留完整的词元历史会使得增加注意力头变得代价高昂。此外,由于注意力机制检索并组合上下文信息,参数缩放也应支持更长的上下文。因此,我们探究注意力参数缩放是否能够直接实现高效且有效的上下文缩放。我们提出了NAMOH,一种架构原生的稀疏注意力机制,每个词元仅激活$H$个注意力头中的$K$个。每个头仅保留其被分配的词元,并在该子序列内执行因果注意力。因此,头的选择共同决定了激活的参数和可用的上下文,而无需扫描完整的历史。在均衡分配下,在固定$K$的同时增加$H$会缩短每个头的历史长度,并减少每个词元的键值(KV)访问量,而不会增加总的KV存储。我们进一步支持头相对旋转位置嵌入,以缩短路由子序列内的位置跨度,旨在减轻由位置引起的注意力噪声。实验表明,NAMOH在总参数相同的情况下可以超越完全激活的模型,同时在激活参数数量匹配的情况下,比更小的稠密模型实现更高效的长上下文推理。它仍然与GQA和现有的稀疏注意力机制兼容。我们希望这项工作能为缩放注意力提供一条新路径,其中参数缩放直接实现上下文缩放。
英文摘要
Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts. We therefore ask whether attention parameter scaling can directly enable efficient and effective context scaling. We introduce NAMOH, an architecture-native sparse attention mechanism that activates $K$ of $H$ heads per token. Each head retains only its assigned tokens and performs causal attention within this subsequence. Head selection thus jointly determines active parameters and available context without scanning the full history. Under balanced assignments, increasing $H$ at fixed $K$ shortens head histories and reduces per-token key-value (KV) access without increasing total KV storage. We further support head-relative rotary position embeddings to shorten positional spans within routed subsequences, aiming to mitigate position-induced attention noise. Experiments show that NAMOH can outperform fully activated models with the same total parameters, while enabling more efficient long-context inference than smaller dense models with matched active parameter counts. It remains compatible with GQA and existing sparse attention mechanisms. We hope this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.