发表机构
University of Science and Technology of China; University College London; Shanghai University of Finance and Economics; AMSS, Chinese Academy of Sciences; University of Copenhagen; Imperial College London; Technical University of Denmark; Shenzhen Research Institute of Big Data(中国科学技术大学; 伦敦大学学院; 上海财经大学; 中国科学院数学与系统科学研究院; 哥本哈根大学; 伦敦帝国学院; 丹麦技术大学; 深圳大数据研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低比特LLM量化中重构损失与实际性能脱节的问题,提出分布鲁棒量化(DRQ)框架,可提升多种PTQ方法量化后模型的下游性能且无推理开销。
AI 中文摘要
仅权重的后训练量化(PTQ)高度依赖重构损失最小化,以在低精度下保持模型质量。我们发现,最小化该损失所偏好的权重,未必能在新任务上带来更好的模型性能;实际上,更低的重构损失甚至会降低模型在相同校准数据上的性能。我们的分析进一步表明,当输入激活分布发生变化时,在校准数据上具有更低重构损失的权重,其损失可能高于其他权重。基于这些观察和分析,我们提出分布鲁棒量化(DRQ),这是一种事后优化流程,可在输入激活分布的受限集合上最小化最坏情况重构损失。DRQ在现有量化网格内优化表示量化权重的整数编码,同时保持量化参数和推理算子不变。大量实验表明,DRQ可提升六种代表性PTQ方法(包括AWQ、GPTQ和ParoQuant)量化后的模型性能,且在稠密和混合专家(MoE)大语言模型上均能取得增益。这些结果确立了DRQ作为仅权重PTQ的通用事后优化框架,无需增加推理开销即可实现更好的下游性能。
英文摘要
Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions. DRQ refines the integer codes representing quantized weights within the existing quantization grid, keeping quantization parameters and inference operators unchanged. Extensive experiments show that DRQ improves models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant, and delivers gains across both dense and mixture-of-experts large language models. These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.