AI 中文总结
本文提出无需训练的ReTopK方法,通过复用历史检索决策加速动态Top-K注意力,在16K-128K上下文下实现低困惑度与高基准分数,128K、K=512时加速3.07倍且仅0.50%困惑度提升。
AI 中文摘要
Top-K稀疏注意力仅关注键值(KV)条目的小子集,从而降低Softmax和值聚合的成本。然而,识别该子集仍需将当前查询与完整KV缓存评分并执行全局Top-K选择,导致选择器成本随上下文长度线性增长,限制了长上下文解码中稀疏注意力的实际效率。本文提出ReTopK,一种无需训练的方法,通过复用历史检索决策加速动态Top-K注意力。ReTopK基于以下观察:相似查询通常关注重叠的支持集,部分重叠的支持集仍可保留精确Top-K注意力的大部分质量。对于每个注意力头,它维护一个有界的历史查询-支持对缓存,为每个新查询检索最相似的缓存查询,将其存储的支持集与最近窗口合并,仅使用当前查询的精确评分对得到的紧凑候选集重新排序。当复用不可靠时,基于相似度的回退机制会调用完整历史的精确Top-K,而定期的精确刷新则限制缓存漂移。ReTopK保留完整KV缓存,仅复用选定的索引,而非历史评分、注意力权重或输出。在16K至128K的上下文范围内,ReTopK在评估的近似方法中实现了最低的PG19困惑度,以及最高的NIAH和LongBench分数。在128K上下文、K=512时,ReTopK与精确Top-K相比仅产生0.50%的困惑度提升,同时将注意力计算加速了3.07倍。
英文摘要
Top-$K$ sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identifying this subset still requires scoring the current query against the full KV cache and performing global Top-$K$ selection, leaving selector cost linear in context length and limiting the practical efficiency of sparse attention for long-context decoding. In this paper, we introduce ReTopK, a training-free method that accelerates dynamic Top-$K$ attention by reusing historical retrieval decisions. ReTopK builds on the observation that similar queries often attend to overlapping supports and that partially overlapping supports can still preserve most of the Exact Top-$K$ attention mass. For each attention head, it maintains a bounded cache of historical query--support pairs, retrieves the most similar cached queries for each new query, unions their stored supports with a recent window, and reranks only the resulting compact candidate set using exact current-query scores. A similarity-based fallback invokes full-history Exact Top-$K$ when reuse is unreliable, while periodic exact refreshes limit cache drift. ReTopK retains the complete KV cache and reuses only selected indices, rather than historical scores, attention weights, or outputs. Across 16K--128K contexts, ReTopK achieves the lowest PG19 perplexity and the highest NIAH and LongBench scores among the evaluated approximate methods. At 128K with $K=512$, ReTopK incurs only a 0.50\% perplexity increase over Exact Top-$K$ while accelerating attention computation by $3.07\times$.
Comments9 pages, 9 figures, and 5 tables