arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoSA:通过代理内核协同设计的稀疏注意力加速长上下文推理

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

Yufei Xue, Lin Niu, Hong Liu, Siran Liu, Hanyong Shao, Wei Liu, Guanghua Yu, Jianchen Zhu, Jun Zhang

arXiv 2607.25291首次发表:更新:

发表机构

Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对自注意力长上下文推理成本高的问题,提出CoSA方法,通过结合内核感知代理与有序跳过内核,在两阶段实现高效稀疏注意力,在主流LLM和长上下文基准测试中,以低预算获高精度,加速注意力并减少端到端时间。

AI 中文摘要

自注意力的二次成本使长上下文推理成本过高,基于代理的块稀疏注意力成为一种实用解决方案。现有方法在预算适中时有效,但预算收紧时,代理会遗漏显著块,内核只能机械应用稀疏掩码,导致模型精度下降。我们提出CoSA,一种代理内核协同设计下的两阶段无训练稀疏注意力方法,它将内核感知代理(KAP)与有序跳过内核(OSK)相结合。在第一阶段,KAP在适中预算下选择块并生成有序掩码,规定内核内循环中访问KV页面的顺序。在第二阶段,OSK应用此掩码并在收紧预算下根据在线softmax统计跳过更多块。在主流LLM主干和长上下文基准测试中,CoSA在较低预算下获得更高精度。令人印象深刻的是,在上下文长度为128K时,CoSA实现了4.93倍的注意力加速,并将端到端首次令牌时间减少了2.53倍,性能下降可忽略不计。

英文摘要

The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical remedy. Existing methods typically rely on a proxy to predict a binary sparse mask and a kernel to consume this mask and perform sparse attention computation. Such an approach is effective under moderate budgets. However, as the budget tightens, the estimated proxy inevitably drops some salient blocks, while the kernel can only apply the sparse mask mechanically, leading to an evident drop in model accuracy. We propose CoSA, a two-stage training-free Sparse Attention under proxy-kernel CO-design, which couples a Kernel-Aware Proxy (KAP) with an Ordered-Skipping Kernel (OSK). In the first stage, the KAP selects blocks under a moderate budget and produces an ordered mask that prescribes the order in which KV pages are visited in the kernel inner loop. In the second stage, the OSK applies this mask and skips more blocks under a tightened budget given online-softmax statistics. Across mainstream LLM backbones and long-context benchmarks, CoSA attains higher accuracy at lower budgets. Impressively, CoSA achieves a 4.93$\times$ attention speedup and reduces end-to-end Time-to-First-Token by 2.53$\times$ under a context length of 128K with negligible performance degradation. Code is available at https://github.com/Tencent/AngelSlim.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑