arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ResidualQuant:循环Transformer的2比特残差KV缓存量化

ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

Heejun Kim, Junyoung Lee, SangLyul Cho, Dongsu Han, Insu Han, Sehoon Kim

arXiv 2610.10381首次发表:更新:

发表机构

KAIST; Yonsei University; Seoul National University(韩国科学技术院; 延世大学; 首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ResidualQuant,利用循环Transformer跨循环KV状态相似性,以最终循环为参考量化残差,结合缩放旋转和混合精度,在INT2下保持精度,减少KV存储80.7%,吞吐量提升显著。

AI 中文摘要

循环Transformer通过多次重复应用共享的Transformer模块来提高参数效率,增加了计算深度而不增加参数数量。然而,KV缓存内存仍然随着循环次数而扩展,成为限制批处理大小和推理吞吐量的关键内存瓶颈。KV缓存量化可以缓解这一瓶颈,但现有方法在激进的低精度设置下常常遭受严重的精度下降。我们观察到循环Transformer提供了一个独特的机会:跨循环的KV状态高度相似。基于这一观察,我们提出了ResidualQuant,它使用最终循环的KV状态作为参考,并用低精度残差表示剩余的循环。我们的方法进一步结合了最小二乘缩放和应用于残差的旋转,以及循环级混合精度,以实现低至INT2的精确量化,同时保持高效的重建。在多个循环Transformer模型以及数学推理和代码生成基准上,ResidualQuant持续改善了与最先进的基于旋转的KV量化相比的精度-内存权衡。特别是,在混合精度设置下,我们的方法保持了接近BF16的精度,同时将理论KV存储减少了80.7%,在相同内存预算下,比基于旋转的基线实现了高达13.0%的精度提升。在RTX 5090上,减少的KV内存流量将固定批处理解码吞吐量提高了高达2.73倍,而更小的内存占用使得批处理大小能够增加高达2倍,峰值吞吐量提高了高达4.15倍。

英文摘要

Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑