arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LiteTopK:利用维度诅咒设计长上下文稀疏注意力中的融合索引器 - TopK 内核

LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention

Ziqi Yin, Jianyang Gao, Peiqi Yin, Jiangneng Li, Gao Cong

arXiv 2607.11976首次发表:更新:

发表机构

Nanyang Technological University; ETH Zurich; The Chinese University of Hong Kong(南洋理工大学; 苏黎世联邦理工学院; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有 GPU 索引器 - TopK 内核效率低的问题,利用维度诅咒设计 LITETOPK 内核,通过采样估计得分范围划分候选结果桶,降低内存开销,在实际部署中加速 GLM 5.2 预填充阶段,提升了效率。

AI 中文摘要

索引器 - TopK 操作在大语言模型的稀疏注意力内核、推荐系统和向量数据库的向量检索中广泛使用。然而,现有的基于 GPU 的索引器 - TopK 内核,如 DeepSeek 稀疏注意力(DSA),由于全局内存流量过大、同步成本高和内存开销大而效率低下。本文利用高维空间中的维度诅咒,设计了一种新颖高效的融合索引器 - TopK 内核 LITETOPK。它先对一小部分数据采样以估计查询 - 数据得分范围,然后利用这些估计在线将候选结果划分为多个桶。这种方式使 LITETOPK 内核能保持紧密的近似阈值,仅回写有希望的候选者,减少不必要的 I/O,大幅降低内存开销,同时保持精确的 Top - k 正确性。实验结果表明,在实际部署场景中,LITETOPK 使 GLM 5.2 的预填充阶段加速 1.2 倍,且内存开销更低。

英文摘要

Indexer-TopK, the operation to compute the scores and select the top-k candidates, is widely used by sparse attention algorithms in large language models and vector retrieval in recommendation systems and vector databases. However, existing GPU-based Indexer-TopK kernels like DeepSeek Sparse Attention (DSA) remain inefficient due to excessive global memory traffic, costly synchronization, and prohibitive memory overhead. In this study, inspired by the curse of dimensionality phenomenon, we first observe that sparse attention scores exhibit a score concentration phenomenon, where scores tend to fall within a narrow range. Based on this observation, we propose LITETOPK, an efficient fused Indexer-TopK kernel. LITETOPK first samples a small subset of data to estimate query-data score ranges, then partitions candidates into bins accordingly. This organization allows the LITETOPK kernel to maintain a tight approximate threshold online, write back only promising candidates, reduce unnecessary I/O and memory overhead while preserving exact Top-k correctness. Building on LITETOPK, we further propose LITEDSA, which exploits the similarity of top-k candidate sets among neighboring tokens. LITEDSA packs neighboring tokens' candidates for joint computation and masks out extra scores for each query, thereby reducing memory traffic while preserving correctness. Experimental results in a real-world deployment environ ment with eight B200 GPUs show that LITETOPK+LITEDSA accelerates the prefill stage of GLM 5.2 by 1.35x, with no performance loss and lower memory overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑