发表机构
Alibaba International Digital Commerce; School of Software Technology, Dalian University of Technology(阿里巴巴国际数字商业集团; 大连理工大学软件学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对长思维链推理模型中标准自注意力计算复杂度高的问题,提出LISA模块,它集成线性注意力模块和闪电索引器,经两阶段训练,能降低推理复杂度,在特定模型上实现推理加速并提升性能。
AI 中文摘要
诸如DeepSeek-R1等长思维链推理模型的最新进展,在测试时缩放范式下实现了越来越长的推理上下文长度。然而,标准自注意力的O(n^2)计算复杂度导致推理成本随长序列急剧增长,限制了长思维链推理在生产环境中的部署。为解决此问题,我们提出了LISA(线性索引稀疏注意力),这是一个即插即用的注意力替换模块,无需从头开始预训练。LISA在原始模型中并行集成了两个轻量级组件:(1)一个线性注意力模块,提供具有O(n)时间复杂度的长距离记忆;(2)一个闪电索引器,从完整上下文中选择前M个重要令牌,输入到稀疏自注意力中。通过门控机制融合两个分支,将生成n个令牌的推理复杂度从O(n^2)降低到O(nM)(M << n)。我们设计了一个两阶段训练管道:阶段1通过集成线性注意力初始化模型以捕获长距离依赖,辅以通过知识蒸馏优化的滑动窗口注意力机制,以近似冻结教师模型的完整自注意力分布。在阶段2中,我们进一步引入索引器以取代静态滑动窗口机制,实现从更广泛上下文中动态选择令牌。索引器使用新颖的逐头KL散度损失进行训练,使其选择行为与教师模型的注意力模式对齐。在DeepSeek蒸馏-Qwen模型上的实验表明,LISA在16K令牌上下文中实现了50%的推理加速,同时在包括AIME和MATH-500在内的推理基准上平均性能提高了5.6%。
英文摘要
Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm. However, the O(n^2) computational complexity of standard self-attention causes inference costs to grow sharply with long sequences, limiting the deployment of long-CoT reasoning in production settings. To address this, we propose LISA (Linear-Indexed Sparse Attention), a plug-and-play attention replacement module that requires no pretraining from scratch. LISA integrates two lightweight components in parallel within the original model: (1) a Linear Attention module that provides long-range memory with O(n) time complexity; (2) a Lightning Indexer that selects the top-M important tokens from the full context to feed into a Sparse Self-Attention. The two branches are fused via a gating mechanism, reducing inference complexity from O(n^2) to O(nM) (M << n) for generating n tokens. We design a two-stage training pipeline: Stage 1 initializes the model by integrating the linear attention to capture long-range dependencies, complemented by a sliding-window attention mechanism that is optimized via knowledge distillation to approximate the full self-attention distribution of a frozen teacher model. In Stage 2, we further introduce the Indexer to replace the static sliding-window mechanism, enabling dynamic token selection from broader contexts. The Indexer is trained using a novel per-head KL divergence loss, which aligns its selection behavior with the attention patterns of the teacher model. Experiments on DeepSeek-distilled-Qwen models demonstrate that LISA achieves a 50% inference speedup under 16K-token context, while improving average performance by 5.6% on reasoning benchmarks including AIME and MATH-500.
Comments20 pages, 10 figures