arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

按需注意力:语言模型知道何时召回

On-Demand Attention: Language Models Know When to Recall

Haibo Feng, Ruiqi Liang, Dongyang Jin, Hanyang Peng, Shiqi Yu

arXiv 2609.20734首次发表:更新:

发表机构

Southern University of Science and Technology; Peking University; Peng Cheng Laboratory(南方科技大学; 北京大学; 鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出按需注意力(ODA),利用轻量级召回头预测全局注意力的益处,选择性调用全局注意力,在保持性能的同时减少全局读取,实现长上下文推理加速。

AI 中文摘要

推理和智能体工作负载日益需要高效的长上下文推理。然而,全注意力解码在每一步都会读取不断增长的历史记录,无论其对下一次预测是否有益。我们证明,预训练模型的解码状态在全局读取之前,已经包含了预测这种益处的信息。基于这一发现,我们引入了按需注意力(ODA),一种局部优先的解码方法,使用轻量级召回头在其预测益处随生成过程变化时,选择性地调用全局注意力。ODA仅训练召回头,保持预训练权重不变,并保留完整的历史KV缓存以供未来召回。我们进一步在vLLM中实现了GPU端的条件执行,将减少的全局读取转化为在长上下文长度下相对于全注意力的实际解码加速。在Qwen和Gemma模型(包括混合注意力骨干)上的实验表明,选择性召回恢复了局部注意力下丢失的大部分性能,同时大幅减少了全局读取。这些发现支持长上下文推理,其中预训练模型引导自身访问其保留的信息。

英文摘要

Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, although the benefit of global access varies across prediction positions. We find that, before global attention is computed for the current step, the decoding states available after local computation in frozen pretrained models already contain information predictive of its benefit over local attention. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method: after local computation, a lightweight recall head decides whether to recompute the current step with global attention. ODA trains only the recall head with modest data and compute budgets, leaving pretrained weights unchanged and retaining the complete historical KV cache so that information skipped at one step remains available for later access. Experiments across model scales and families, including hybrid attention backbones, show that ODA recovers most of the performance lost under local attention while substantially reducing the frequency of global attention. Controlled long-context measurements in vLLM further show that GPU-side conditional execution translates fewer global reads into practical decoding speedups over full attention. These findings show that pretrained decoding states can support both token prediction and decisions about accessing distant information, allowing models to allocate global computation as needed during decoding.

Comments30 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑