arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08450cs.DC

样本引导的长上下文稀疏注意力精确 Top-K 选择

Sample-Guided Exact Top-K Selection for Long-Context Sparse Attention

发表机构腾讯公司
查看机构详情
  • Tencent Inc.(腾讯公司)

机构由 AI 辅助整理,请以论文原文为准。

Siran Liu, Yang Xue, Theo Tang, Changxu Shao, Qian Cheng, Haimeng Ren, Donghua Jiang, Haipeng Ming, Lehua Ding, Zhonghan Lin, Shengying Wei, Wei Liu, Kai Liu, Jianchen Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

提出样本引导的精确 Top-K 选择方法,利用固定步长部分视图估计边界,减少长上下文稀疏注意力计算开销,在多种配置下实现最高 1.75 倍加速。

中文摘要 AI 辅助

稀疏注意力通过保留固定大小的索引令牌子集来限制下游注意力计算,但其独立的精确 Top-K 阶段仍必须处理随上下文长度增长的材料化得分行。生产级基数选择器仅在完整行遍历后才能发现第一个可操作的边界,迫使在精确细化之前进行另一次行级遍历。我们观察到,定位紧凑的上尾所需的精度远低于识别精确秩边界,并且当前行的固定步长部分视图在不同长度下仍与相应的完整行秩保持校准。我们提出了 HPC-Ops Top-K,一种用于不规则稀疏注意力得分行的样本引导精确选择器。固定步长视图提出行局部粗略边界;强制性的完整行遍历验证其充分性,形成被接受的候选集,并在未解决的前沿上初始化精确 FP32 细化。嵌套的次级边界和精确恢复在输出提交前处理填充不足的提议,因此采样控制常见路径工作但绝不牺牲正确性。GPU 实现融合了完整行验证和候选形成,并在可图捕获的不规则行调度背后结合了持久、KV 分割和直接精确执行。我们在 Hy4-Preview 的索引器得分上评估了 HPC-Ops Top-K。在 20 种算子配置中,它比最快的已验证外部精确基线快 1.29--1.75 倍,几何平均加速比为 1.55 倍。它还在两个框架派生的稀疏注意力轨迹上实现了 1.36 倍和 1.48 倍的加速。该实现可在腾讯开源的高性能 LLM 推理算子库 HPC-Ops 中获取,网址为 https://github.com/Tencent/HPC-Ops。

英文摘要

Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top-$K$ stage must still process materialized score rows whose length grows with context. Production radix selectors discover their first actionable boundary only after a complete-row pass, forcing another row-scale traversal before exact refinement. We observe that locating a compact upper tail requires substantially less resolution than identifying the exact rank boundary, and that fixed-stride partial views of the current row remain calibrated to the corresponding complete-row rank across ragged lengths. We present HPC-Ops Top-K, a sample-guided exact selector for ragged sparse-attention score rows. A fixed-stride view proposes a row-local coarse boundary; the mandatory complete-row pass certifies its sufficiency, forms the admitted candidate set, and initializes exact FP32 refinement over the unresolved frontier. A nested secondary boundary and exact recovery handle underfilled proposals before any output is committed, so sampling controls common-path work but never correctness. The GPU implementation fuses complete-row certification and candidate formation, and combines persistent, KV-split, and direct-exact execution behind graph-capturable ragged-row dispatch. We evaluate HPC-Ops Top-K on indexer scores from Hy4-Preview. It outperforms the fastest verified external exact baseline by $1.29$--$1.75\times$ across 20 operator configurations, with a $1.55\times$ geometric-mean speedup. It further achieves $1.36\times$ and $1.48\times$ speedups on two framework-derived sparse-attention traces. The implementation is available in HPC-Ops, Tencent's open-source high-performance operator library for LLM inference, at https://github.com/Tencent/hpc-ops.

补充信息

↑