arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20971cs.AI

RBS-Attention:面向长上下文大语言模型的半径受限稀疏预填充

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

Chuxu Song, Jiuqi Wei, Zhencan Peng

首次发表
浏览论文内容

中文总结 AI 辅助

针对长上下文预填充中块质心隐藏相关令牌的问题,提出RBS-Attention,通过质心基础分支和基于最大键块半径的救援分支进行双分支稀疏选择,在保持质量的同时实现显著加速。

中文摘要 AI 辅助

长上下文大语言模型推理日益受到预填充阶段的限制,在此阶段,密集自注意力在生成开始前处理整个提示。稀疏块选择可以降低这一成本,但块质心可能将高度相关的令牌隐藏在众多无关令牌之中。我们将这种失效模式称为均值稀释,并提出RBS-Attention,一种无需训练的稀疏预填充方法,具有两个互补的选择分支。质心基础分支捕获平均相关性,而救援分支利用最大键块半径及其随提示、层和头变化的分布来识别存在低估风险的块。对两个分支独立设置阈值并组合其掩码,可控制救援块的贡献,同时保持常规块稀疏FlashAttention的执行。在H100 GPU上,RBS-Attention在Qwen3-30B-A3B-Instruct-2507-FP8的128K上下文下实现了20.65倍的独立预填充注意力加速、11.92倍的vLLM预填充注意力加速和5.97倍的端到端首令牌时间加速。在密集Qwen3-32B模型上,其RULER总体准确率为88.65,而密集注意力为89.52;LongBench-v2、InfiniteBench和Video-MME提供了额外的质量评估。支持性实验测量了实际保留率,在匹配密度下比较了选择器,并表征了块大小、阈值和内存行为。这些结果共同支持半径自适应双分支选择作为长上下文预填充的有效方法。

英文摘要

Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution to identify blocks at risk of underestimation. Independently thresholding the two branches and combining their masks controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution. On H100 GPUs, RBS-Attention achieves 20.65$\times$ standalone prefill-attention speedup, 11.92$\times$ vLLM prefill-attention speedup, and 5.97$\times$ end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8. On the dense Qwen3-32B model, it obtains 88.65 overall RULER accuracy versus 89.52 for dense attention; LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation. Supporting experiments measure actual retention, compare selectors at matched density, and characterize block-size, threshold, and memory behavior. Together, these results support radius-adaptive dual-branch selection as an effective approach to long-context prefill.

发表机构

  • Rutgers University(罗格斯大学)
  • OceanBase, Ant Group(OceanBase,蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

↑