arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38169cs.CLcs.AIcs.LG

STEPQuant:Delta规则循环状态量化中误差何时何地重要

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei

首次发表
浏览论文内容

中文总结 AI 辅助

针对Delta规则循环状态量化误差的时间与空间影响,提出STEPQuant后训练量化框架,按误差幅度和记忆寿命分配精度,在6位预算下匹配FP32精度,并实现超5倍状态压缩及高达68.7%的内存减少。

中文摘要 AI 辅助

线性注意力用固定大小的循环状态替代不断增长的KV缓存,然而在并发服务下,这些持久状态可能成为显著的内存瓶颈。直接将循环状态量化为低精度往往会导致严重的精度下降,因为量化误差会通过连续的状态更新而传播。我们发现这些误差的影响取决于两个互补的维度:在时间上,长寿命记忆中的误差可能跨多个解码步骤持续存在;在空间上,不同键行中的误差对模型输出的影响不同,同时状态幅度在行和列方向上均有显著变化。基于这些观察,我们提出了STEPQuant,一个针对Delta规则循环状态的空间-时间后训练量化框架。STEPQuant根据误差幅度和记忆寿命分配精度,并基于状态分布和键行对输出误差的影响联合拟合键行和值列的缩放因子。在Qwen3.8-27B和Kimi-Linear-48B-A3B-Instruct上跨越长生成和短生成基准的实验表明,STEPQuant在标称6位预算下紧密匹配FP32状态的精度,并在其4位配置中优于均匀INT8。集成到带有优化GPU内核的SGLang中,6位STEPQuant实现了超过5倍的循环状态压缩,并将总服务内存减少高达68.7%。我们的代码可在该https URL获取。

英文摘要

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.

发表机构

  • City University of Hong Kong(香港城市大学)
  • Harvard University(哈佛大学)
  • Zhejiang University(浙江大学)
  • NLPR & MAIS, Institute of Automation, CAS(中国科学院自动化研究所模式识别国家重点实验室与多模态人工智能系统全国重点实验室)
  • The Chinese University of Hong Kong(香港中文大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑