AI 中文总结
研究针对令牌级稀疏注意力中索引器瓶颈问题,提出PIVOT方法,通过共享前缀扫描为附近查询高效选前k个令牌,有两个速度与保真度权衡的变体,在多个模型上实验,能匹配密集索引器准确性,大幅加速并降低长上下文延迟。
AI 中文摘要
在生产系统中,由DeepSeek稀疏注意力(DSA)实现的令牌级稀疏注意力使下游注意力高效,但将瓶颈转移到了为其提供数据的索引器上。为为每个查询选择前k个令牌,索引器仍需对每个先前令牌进行评分,对于长度为L的序列,每层成本为O(L^2)。研究发现每个查询扫描在很大程度上是冗余的。PIVOT(通过一次完整前缀遍历的代理索引)利用这些属性,它是一种无需训练、可直接替代DSA索引器的方法,能在一组附近查询中共享一次前缀扫描。PIVOT将一组聚合为单个代理查询,进行一次共享的完整前缀扫描以获得候选集,然后为每个查询从该集中选择前k个。有两个变体在速度和保真度之间进行权衡。在DeepSeek-V3.2和GLM-5.1上,PIVOT在长上下文时匹配密集DSA索引器的准确性,同时加速高达4倍,端到端延迟降低高达1.6倍。
英文摘要
Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2) per layer for a sequence of length L. We observe that this per-query scan is largely redundant: nearby queries select highly overlapping top-k tokens, and the indexer scores are long-tailed along the key axis. We exploit these properties in PIVOT, Proxy Indexing Via One full-prefix Traversal, a training-free, drop-in replacement for the DSA indexer that shares one prefix scan across a group of nearby queries. PIVOT aggregates a group into a single proxy query, performs one shared full-prefix scan to obtain a candidate set, and then selects a top-k for each query from that set. Two variants trade speed for fidelity: PIVOT-Reuse shares the proxy top-k across the group for maximum speed, whereas PIVOT-Refine re-scores the candidate set with the indexer of each query and then selects an individual top-k, matching the dense indexer at a small additional cost. A single algorithm covers both inference phases, differing only in how groups are formed: fixed-size groups of consecutive queries in prefill, and the queries decoded together in one multi-token prediction (MTP) step in decode. On DeepSeek-V3.2 and GLM-5.1 across LongBench and RULER, PIVOT matches the accuracy of the dense DSA indexer while accelerating it by up to 4x and reducing end-to-end latency by up to 1.6x at long context.