PageWeaver:面向稀疏注意力的KV引导查询分组
PageWeaver: KV-Guided Query Unions for Sparse Attention
浏览论文内容
中文总结 AI 辅助
PageWeaver是一种KV引导的稀疏注意力查询分组执行设计,通过GPU搜索和双CTA内核实现,在H200等硬件上提升了预填充吞吐量与调用速度,分离了执行组复用的相关成本。
中文摘要 AI 辅助
动态稀疏注意力会限制每个查询选择的KV页,但有限的支持集并不一定能带来高效的GPU计算。查询分组可共享页加载并填充Tensor Core块,其成本取决于将哪些查询分组在一起。我们提出PageWeaver,这是一种利用所选页的亲和性来组装查询组的执行设计,同时保留每个查询的原始支持集和完整输出所有权。有界GPU搜索会生成查询ID,而感知ID的双CTA内核会处理这些查询ID,无需实例化重新排序的Q张量或跨页部分输出。直接的KV页分组实现为非局部复用和归约成本提供了补充性设计研究。在全程使用FP8 KV的情况下,H200 Union8实现在6个捕获样本上实现了比实测FlashInfer路径高1.70倍的几何平均完整调用加速。在线重新分组在5个选定的64K上下文捕获上将延迟进一步降低了3.26-7.66%。全模型预填充吞吐量比测试的原生路径高7.88-14.36%;增量重新分组收益较小,在32K/64K时观察到的中位数增益为0.47-0.73%,在8K时出现性能下降。B300对比确定了准备成本和更强的原生内核会消除该方法优势的情况。这些结果将执行组复用与在线利用它的完整成本分离开来。
英文摘要
Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership. A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs. A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost. With FP8 KV throughout, the H200 Union8 implementation achieves a 1.70x geometric-mean complete-call speedup over the measured FlashInfer path on six captures. Online regrouping further lowers latency by 3.26-7.66% on five selected 64K-context captures. Whole-model prefill throughput is 7.88-14.36% above the tested native path; the incremental regrouping benefit is smaller, with observed median gains of 0.47-0.73% at 32K/64K and regressions at 8K. A B300 comparison identifies cases where preparation cost and a stronger native kernel remove the advantage. These results separate execution-group reuse from the complete cost of exploiting it online.
发表机构
- SenseTime(商汤科技)
机构由 AI 辅助整理,请以论文原文为准。