arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27513cs.LGcs.AI

DAMP:感知衰减的混合精度循环状态量化

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

  • South China University of Technology(华南理工大学)
  • Meituan(美团)
  • East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zunhai Su, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng

AI总结:

DAMP是针对GDN、KDA类语言模型循环状态的混合精度量化方法,通过识别高风险通道差异化精度存储,在接近FP32基线精度的同时大幅降低存储、加速更新并减少推理延迟。

AI中文摘要:

Softmax注意力会为每个前序token存储键(Key)和值(Value)向量,导致推理内存随序列长度增长。近期结合门控DeltaNet(GDN)或Kimi Delta注意力(KDA)的语言模型,通过用固定大小的循环状态替换多数层的KV缓存降低了该成本,但这些循环状态通常以FP32存储,会消耗大量GPU内存;其更新受内存带宽限制,是解码延迟的重要来源。据我们所知,我们是首个研究基于GDN和KDA的语言模型中循环状态的训练后量化的团队。我们发现均匀量化的精度-存储权衡效果差:INT8和FP8已在复杂推理任务上降低精度,而INT4和NVFP4会将精度降至接近零。我们还发现,大部分量化误差能量集中在一小部分通道中,且状态通道的相对衰减强度在不同提示和任务间保持稳定。基于这些发现,DAMP在离线校准期间同时利用量化误差能量和基于衰减的持久性来识别高风险通道,将这些通道以更高精度存储,其余通道以INT8存储。我们在Qwen3.6-35B和Kimi-Linear-48B上,针对涵盖数学推理、通用推理和代码生成的六个基准评估了DAMP。在每个状态值9.9位的情况下,DAMP保持的平均精度接近FP32基线,将循环状态存储降低69.1%,使循环状态更新内核加速最高达2.01倍,并降低全模型的每输出词元时间(TPOT)最高达10.9%。

英文摘要:

Complex reasoning and agentic applications increasingly rely on long-context inference, where growing KV caches increase both memory usage and decoding overhead. Hybrid models reduce these costs by combining Softmax Attention with Gated DeltaNet (GDN) or Kimi Delta Attention (KDA), which maintain fixed-size recurrent states. These states are commonly stored in FP32 and consume substantial GPU memory, while their updates are limited by memory bandwidth. Quantization can reduce both storage footprint and memory traffic, but we find that uniform INT8 and FP8 degrade complex reasoning accuracy, while INT4 and NVFP4 collapse it to near zero. To our knowledge, this is the first study of post-training recurrent-state quantization for GDN and KDA. Our analysis reveals that outliers in GDN and KDA states are concentrated in particular key channels and value dimensions. Learned decay influences how much quantization error is retained. We find that largely the same GDN heads and KDA key channels exhibit slow decay across tasks. Based on these insights, we propose DAMP, which jointly considers quantization error and decay-based error retention to select high-risk key channels offline. Under a fixed storage budget, it retains these channels in FP16 and stores the remainder in INT8. We evaluate DAMP on Qwen3.6-35B, Kimi-Linear-48B and Kimi-K3 across six reasoning and code generation benchmarks. At 9.9 bits per state value, DAMP maintains average accuracy close to FP32. In SGLang, DAMP reduces recurrent-state storage by 69.1%, accelerates the recurrent-state update kernel by up to 2.59x , and lowers full-model time per output token by up to 19.0%.

↑