LoGo:令牌级动态局部-全局注意力
LoGo: Token-Level Dynamic Local-Global Attention
浏览论文内容
中文总结 AI 辅助
针对Transformer长上下文注意力的计算瓶颈,提出令牌级动态局部-全局注意力LoGo,通过学习门控动态分配注意力范围,提升长距离检索性能并实现计算加速。
中文摘要 AI 辅助
随着上下文长度的增加,注意力机制逐渐成为大型语言模型的主要计算瓶颈。标准Transformer功能强大但计算效率低下,因为它会为每个令牌分配相同的注意力预算,而不考虑其上下文需求。现有的局部-全局混合模型通过结合受限上下文注意力和全上下文注意力,提供了更高效的替代方案,但它们通常在各层或各头部静态分配注意力范围。为解决这些局限性,我们提出了LoGo,一种令牌级动态局部-全局注意力机制,它将注意力范围作为注意力预算分配的直接代理。每个LoGo层包含耦合的局部和全局分支:所有令牌都在受限上下文窗口上接收高效的局部注意力,而一个学习到的门控仅为需要长距离信息的令牌激活具有全上下文访问权限的全局注意力。基于阈值的预算控制器在无辅助损失的情况下维持目标全局比例,渐进式掩码调度在稀疏路由生效前稳定训练。我们进一步实现了查询稀疏的Triton内核,将缩减后的全局注意力计算转化为实际的加速效果。大量实验验证了LoGo的有效性,表明它在不同模型规模下都保留了全注意力Transformer的缩放特性。在受控对比中,LoGo优于全注意力Transformer和匹配预算的静态局部-全局混合模型,在长距离检索任务上取得了明显提升。分析进一步显示,LoGo学习到了可解释的范围分配模式。这些结果表明,学习到的令牌级范围分配是改善长上下文场景下性能-计算权衡的有效且可扩展的方法。
英文摘要
As context lengths scale, attention increasingly becomes a primary computational bottleneck in large language models. Standard Transformers remain powerful but computationally inefficient, as they allocate the same attention budget to every token regardless of its contextual demand. Existing local-global hybrids provide a more efficient alternative by mixing restricted- and full-context attention, but they typically allocate span statically across layers or heads. To address these limitations, we propose LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation. Each LoGo layer contains coupled local and global branches: all tokens receive efficient local attention over a restricted context window, while a learned gate activates global attention with full-context access only for tokens requiring long-range information. A threshold-based budget controller maintains a target global ratio without auxiliary losses, and a progressive masking schedule stabilizes training before sparse routing takes effect. We further implement query-sparse Triton kernels that convert reduced global-attention computation into practical speedups. Extensive experiments validate LoGo's effectiveness, showing that it preserves the scaling behavior of full-attention Transformers across model sizes. In controlled comparisons, LoGo improves over the full-attention Transformer and matched-budget static local-global hybrids, with clear gains on long-range retrieval. Analysis further shows that LoGo learns interpretable span allocation patterns. These results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。