arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意力路由早期稳定:循环语言模型的工作集推理

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

Ke Wan, Chen Chen

arXiv 2609.27373首次发表:更新:

发表机构

University of Virginia(弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该论文提出WISE方法,利用循环语言模型中注意力早期稳定的特性,通过复用稀疏工作集支持来减少重复计算,在保持多跳问答性能的同时实现高达1.76倍的注意力加速。

AI 中文摘要

循环语言模型反复应用共享网络块来细化潜在表示,但标准推理在每个循环步骤都会重新计算全局注意力。我们研究了跨循环深度的注意力动态,发现注意力支持和分布在隐藏状态及注意力输出之前就显著稳定下来。这表明存在一个两阶段结构:早期步骤发现相关上下文的稀疏工作集,而后期步骤则在基本相同的路由支持上细化表示。受此结构启发,我们提出了WISE(工作集推理与支持利用),一种无需训练的方法,在早期循环中使用不受限制的全局注意力,后期直接复用已发现的块结构支持,同时保持循环深度和支撑内注意力计算的动态性。受控干预表明,循环发现很重要,仅支持复用比更严格的注意力复用替代方案更能保持模型行为。在多跳问答基准测试中,WISE在很大程度上保持了全注意力性能,而上下文扩展则揭示了越来越稀疏的工作集和更高的效率提升。在2K上下文内质量基本保持,在4K时出现可测量的损失。优化的稀疏注意力实现在4K时比原生FlashAttention实现了高达1.76倍的注意力加速,在完整的32步注意力轨迹上实现了1.36倍的加速。我们的代码可在https://github.com/tbn5pj/WISE_code获取。

英文摘要

Recurrent-depth language models, such as looped Transformers, repeatedly apply shared network blocks to refine latent representations without generating explicit intermediate reasoning tokens. However, each step recomputes full attention over the entire context, repeating costly global routing. We study how attention routing evolves across recurrent depth and find a consistent separation in convergence timescales: attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests two stages of recurrent inference: early discovery of a sparse working set, followed by representation refinement over largely stable routing support. Motivated by this finding, we introduce WISE (Working-set Inference with Support Exploitation), a training-free method that uses unrestricted attention during early recurrent steps to discover a block-structured working set, then reuses its support in later steps while keeping attention weights and recurrent refinement dynamic. Controlled interventions show that multi-step discovery yields more effective working sets than first-step selection, and that support reuse better preserves model behavior than more restrictive forms of attention reuse. Across multi-hop QA benchmarks, WISE largely preserves full-attention performance. Matched context-scaling experiments reveal an increasingly favorable quality-efficiency tradeoff as routing support becomes sparser with longer contexts. A sparse-attention implementation achieves up to a 1.76x late-step attention speedup over native FlashAttention at 4K context. Code: https://github.com/tbn5pj/WISE_code.

CommentsCode: https://github.com/tbn5pj/WISE_code

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑