QUILT:通过共享查询执行重新思考稀疏注意力预填充
QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
浏览论文内容
中文总结 AI 辅助
QUILT是一种工作负载感知的稀疏注意力执行机制,通过共享查询的KV条目减少冗余开销,在LongBench上的实验显示其可降低内核延迟、KV数据量及TTFT延迟,且精度损失极小。
中文摘要 AI 辅助
稀疏注意力可降低长上下文注意力的成本,但现有内核通常独立处理查询,重复加载和反量化跨查询共享的KV条目。我们观察到相邻查询选择的KV条目存在大量重叠,这为跨查询复用创造了机会。我们提出QUILT,一种工作负载感知的稀疏注意力执行机制,它联合处理相邻查询并复用共享KV条目,以减少冗余内存流量和计算。QUILT引入了移位-比较集分解(Shift-and-Compare Set Decomposition,SCSD),该方法将不规则集操作转换为适合现代加速器的规则数据并行原语,并将SCSD与注意力计算流水线以隐藏其开销。级联共享在多个粒度上分层捕获复用。感知块的执行策略在共享粒度与硬件块利用率之间取得平衡,并选择性移除低重要性的查询特定尾部,以消除利用率不足的块。我们在LongBench上使用GLM-5.3和DeepSeek-3.2,在张量并行和序列并行两种设置下评估QUILT。与最先进的稀疏注意力内核相比,QUILT将平均内核延迟降低了高达55.1%,处理的KV数据减少了高达55.9%,同时将首token生成时间(TTFT)延迟降低了高达36.8%,且精度下降可忽略不计。
英文摘要
Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries. We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse. We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation. QUILT introduces Shift-and-Compare Set Decomposition (SCSD), which transforms irregular set operations into regular data-parallel primitives suitable for modern accelerators, and pipelines SCSD with attention computation to hide its overhead. Cascaded sharing captures reuse hierarchically at multiple granularities. A tile-aware execution strategy balances sharing granularity with hardware tile utilization and selectively removes low-importance query-specific tails to eliminate underutilized tiles. We evaluate QUILT on LongBench using GLM-5.3 and DeepSeek-3.2 under both tensor and sequence parallelism. Compared with the state-of-the-art sparse-attention kernel, QUILT reduces average kernel latency by up to 55.1% and processed KV data by up to 55.9%, while reducing time-to-first-token (TTFT) latency by up to 36.8% with negligible accuracy degradation.
发表机构
- Huawei Technologies Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。