arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38166cs.LGcs.AI

LeapQuant:具有精确循环状态量化机制的高效线性注意力

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

发表机构加州大学伯克利分校 · 华盛顿大学 · 麻省理工学院
另 2 家 · 查看机构详情
  • UC Berkeley(加州大学伯克利分校)
  • University of Washington(华盛顿大学)
  • MIT(麻省理工学院)
  • Perplexity AI
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica

首次发表
浏览论文内容

中文总结 AI 辅助

针对线性注意力循环状态量化误差累积与离群值问题,提出免训练方法LeapQuant,通过逐窗口量化与补偿token,在8位量化下实现近乎无损性能并显著加速推理。

中文摘要 AI 辅助

最近的LLM越来越多地采用混合设计,用线性注意力取代标准注意力,例如Gated DeltaNet (GDN)和Kimi Delta Attention (KDA)。尽管这些设计将上下文压缩为固定大小的循环状态,并大幅降低了长上下文处理的成本,但反复读取和更新该状态仍是主要的推理瓶颈。量化提供了一种自然的降低该成本的方法,但可能由于舍入误差的累积以及状态中离群行和列的存在而显著降低模型质量。为应对这些挑战,我们提出LeapQuant,一种免训练方法,在8位循环状态量化下实现近乎无损的性能。首先,为缓解误差累积,我们提出逐窗口量化,它跳过一段窗口的token,并仅在其末尾对状态进行一次量化。在窗口内,输出由固定的低位状态与高精度缓冲更新共同计算。其次,为减少每次量化引入的误差,LeapQuant将状态中最大的离群值保留为少量高精度补偿token,这些token共享真实token的更新路径。然后,我们在量化前对剩余残差进行平滑处理以进一步减少误差。在Qwen、Kimi和GLM模型系列上的综合实验表明,LeapQuant在推理过程中显著降低了内存和计算成本。在精度与FP32基线相当的情况下,它在NVIDIA B200、RTX PRO 6000和RTX 5090 GPU上实现了内核级别平均2.05--3.70倍的加速,以及端到端推理1.47倍的加速。

英文摘要

Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.

补充信息

相关深度报道

↑