发表机构
Theory Lab, 2012 Labs, Huawei Technologies Co., Ltd.(华为技术有限公司,2012实验室,理论实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对动态稀疏注意力KV缓存卸载中的内存瓶颈,提出AVSG算子,通过向量化哈希匹配、均衡H2D传输和基于生命周期的缓冲区管理,实现高效重用,显著提升吞吐并降低延迟。
AI 中文摘要
动态稀疏注意力通过仅选择一部分token来减少长上下文注意力计算,但仍需访问完整的KV缓存,导致服务受内存带宽限制。将KV缓存卸载到主机内存可减轻设备内存压力,但将H2D传输置于解码关键路径上。在DSA中,不同解码步骤间所选KV条目的大量重叠为设备驻留重用创造了机会。然而,高效利用这种重用面临三个关键挑战:所选token与HBM缓冲区中保留条目匹配成本高昂、剩余H2D传输在加速器核心间分布不均、以及如何在有限HBM容量内保留频繁访问的KV条目。我们提出AVSG,一个针对DSA解决这些挑战并保持其精确token选择的算子。AVSG使用向量化哈希匹配来识别可重用的HBM缓冲区槽位,并为需要传输的条目预留槽位,然后将这些H2D传输均匀分配到加速器核心。基于生命周期的缓冲区管理在解码步骤间将频繁访问的KV条目保留在设备缓冲区中,共享槽位预留将重用扩展到多token预测迭代内的token间。在单个NPU上,向量化哈希匹配比标量双指针匹配快2.80倍,仅缺失传输相比请求级分配将有效H2D带宽提升高达34.44倍,8K条目缓冲区在真实请求上达到94.83%的命中率。这些增益转化为端到端改进:在处理真实生产数据集的服务栈上,与无HBM缓冲区重用的相同KV卸载布局相比,AVSG将每个输出token的时间减少39%,输出吞吐量提升1.27倍,展示了高效匹配、均衡传输和有效HBM驻留在生产规模服务中的益处。
英文摘要
Dynamic sparse attention reduces long-context attention computation by selecting only a subset of tokens, but still requires access to the full KV cache, leaving serving memory-bound. Offloading the KV cache to host memory reduces device memory pressure but places H2D transfers on the decoding critical path. In DSA, substantial overlap in selected KV entries across decoding steps creates an opportunity for device-resident reuse. Exploiting this reuse efficiently, however, presents three critical challenges: costly matching of selected tokens against entries retained in the HBM buffer, uneven distribution of the remaining H2D transfers across accelerator cores, and retaining frequently accessed KV entries within limited HBM capacity. We present AVSG, an operator that addresses these challenges for DSA while preserving its exact token selections. AVSG uses vectorized hash matching to identify reusable HBM-buffer slots and reserve slots for entries requiring transfer, then evenly partitions these H2D transfers across accelerator cores. Lifetime-based buffer management retains frequently accessed KV entries in the device buffer across decoding steps, and shared slot reservations extend reuse across tokens within a multi-token prediction iteration. On a single NPU, vectorized hash matching is 2.80x faster than scalar dual-pointer matching, miss-only transfer raises effective H2D bandwidth by up to 34.44x over request-level assignment, and an 8K-entry buffer reaches a 94.83% hit rate on real requests. These gains translate into end-to-end improvement: on a serving stack processing a real-world production dataset, AVSG reduces time per output token by 39% and increases output throughput by 1.27x relative to the same KV offload layout without HBM-buffer reuse, demonstrating the benefit of efficient matching, balanced transfer, and effective HBM residency in production-scale serving.
Comments20 pages, 7 figures