JustQuant:4比特激活量化无需平滑、SVD或旋转
JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
浏览论文内容
中文总结 AI 辅助
JustQuant提出仅用普通低比特算子实现4比特激活量化,通过多层级蒸馏方法Theseus QAD将复杂度转移到训练中,提升量化质量并避免部署时的复杂算子。
中文摘要 AI 辅助
近年来,生成模型变得日益强大,但其推理成本持续增长。模型量化提供了一种有前景的方法来压缩这些模型并加速推理。然而,在4比特精度下,激活量化比权重量化更具挑战性。最近的训练后量化(PTQ)和量化感知训练(QAT)方法通过引入平滑、SVD分支、旋转、混合精度或高级格式(如NVFP4)在4比特激活量化方面取得了进展。这些额外的算子和数据类型对推理引擎和硬件提出了苛刻要求,限制了低精度模型的广泛采用。能否仅使用普通的低比特算子实现量化?为了回答这个问题,我们提出了JustQuant,一个简单而有效的框架,将低比特量化的复杂性从部署时的算子转移到训练过程中。我们首先从知识蒸馏的角度重新审视模型量化,并表明现有PTQ和QAT方法失败的一个关键原因是它们通常仅在单一层级利用监督信号。然后,我们引入了Theseus QAD,一种量化感知蒸馏方法,逐步应用多层级监督,类似于忒修斯之船中的渐进替换过程。在DiT和扩散大语言模型上的大量实验显示了两种不同的机制。对于较小的模型,Theseus QAD可以作为轻量级的预热阶段,显著提升后续使用普通算子的QAT性能,而朴素的QAD在同一设置下可能崩溃。对于较大的模型,Theseus QAD提供了比普通QAD更强的蒸馏训练路径。在两种机制下,JustQuant都提高了低比特量化的质量,同时避免了现有许多PTQ方法所需的复杂算子。
英文摘要
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-training quantization (PTQ) and quantization-aware training (QAT) methods have made progress in 4-bit activation quantization by introducing smoothing, SVD branches, rotations, mixed precision, or advanced formats such as NVFP4. These additional operators and data types impose demanding requirements on inference engines and hardware, limiting the broad adoption of low-precision models. Can quantization be achieved using only plain low-bit operators? To answer this question, we propose JustQuant, a simple yet effective framework that moves the complexity of low-bit quantization from deployment-time operators into the training process. We first revisit model quantization from the perspective of knowledge distillation and show that a key reason existing PTQ and QAT methods fail is that they typically exploit supervision at only a single level. We then introduce Theseus QAD, a quantization-aware distillation method that progressively applies multi-level supervision, analogous to the gradual replacement process in the Ship of Theseus. Extensive experiments on DiT and diffusion large language models show two distinct regimes. For smaller models, Theseus QAD can serve as a lightweight warm-up stage that substantially improves subsequent QAT with plain operators, while naive QAD may collapse in the same setting. For larger models, Theseus QAD provides a stronger distillation training path than ordinary QAD. Across both regimes, JustQuant improves low-bit quantization quality while avoiding the complex operators required by many existing PTQ methods.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Tsinghua University(清华大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。