SPIN:用于稀疏注意力的影子预测索引器
SPIN: Shadow Predictive Indexer for Sparse Attention
浏览论文内容
中文总结 AI 辅助
SPIN通过轻量级历史预测识别重要KV块,避免全缓存评分,在长上下文和智能体任务中实现30-40%稀疏度,提升vLLM吞吐量14.9%并降低延迟13.2%。
中文摘要 AI 辅助
基于索引器的稀疏注意力通过仅将固定数量且少量的重要令牌传递给核心注意力,从而降低其成本。然而,索引器必须在每个解码步骤对整个KV缓存进行评分。随着上下文长度的增长,这种评分开销成为主要瓶颈。我们提出SPIN(影子预测索引器)来减少这种索引器开销。SPIN使用轻量级的、基于历史的预测来识别重要的KV块,避免了在每个解码步骤对整个KV缓存进行评分的需要。SPIN将KV块和推测解码视为一等设计和实现考虑因素。在长上下文和智能体基准的广泛评估中,SPIN实现了30-40%的稀疏度,同时保持了任务质量。在端到端的vLLM服务中,SPIN将输出吞吐量提高了最多14.9%,并将中位令牌间延迟降低了最多13.2%。
英文摘要
Indexer-based sparse attention reduces the cost of core attention by passing only a fixed, small number of important tokens to it. However, the indexer must still score the entire KV cache at every decoding step. This scoring overhead becomes a major bottleneck as the context length grows. We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead. SPIN uses lightweight, history-based prediction to identify important KV blocks, avoiding the need to score the full KV cache at every decoding step. SPIN treats KV blocks and speculative decoding as first-class design and implementation considerations. Across extensive evaluations on long-context and agentic benchmarks, SPIN achieves 30-40% sparsity while preserving task quality. In end-to-end vLLM serving, SPIN improves output throughput by up to 14.9% and reduces median inter-token latency by up to 13.2%.