arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20953cs.CLcs.AIcs.LGcs.PF

量化感知修复:恢复压缩4位大语言模型的实用方案

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús

首次发表
浏览论文内容

中文总结 AI 辅助

针对压缩4位大模型部署时性能下降问题,提出量化感知修复方案,直接从原始未压缩模型蒸馏4位学生模型,在多基准测试中表现优于量化感知训练,且部署便捷。

中文摘要 AI 辅助

以低成本部署大语言模型,越来越意味着要交付同时完成结构压缩(参数仅为原模型的一小部分)和4位量化的模型。这两个步骤会严重损害模型的推理、数学运算、编码和长上下文处理能力,因此在部署前需要一个恢复(或修复)阶段。默认的方案是量化感知训练(Quantization-Aware Training, QAT),该方案将压缩后的量化模型重新拟合到硬标签;在我们的流程中,它收敛缓慢且会在达到峰值后崩溃。我们转而采用量化感知修复(Quantization-Aware Healing, QAH)。由于结构压缩后的模型从未以全精度独立训练过,其bfloat16检查点是通过蒸馏恢复的原模型近似值;QAH则直接从原始未压缩模型中蒸馏出4位学生模型。在GPT-OSS 120B→60B→MXFP4的流程中,QAH学生模型在9个基准测试中的7个上表现与bfloat16源模型相当或更优,权重内存约减少4倍,参数数量仅为教师模型的一半,且以开放权重形式发布,命名为Hypernova-60B。与匹配的QAT基线相比,它达到可比峰值的速度快约7倍,在持续训练下保持稳定,无需手动调整的早停策略。我们还报告了部署相关的经验教训,包括分布式训练后端之间存在的巨大且可复现的质量差距。我们的目标是提供一个无需进行数周超参数搜索即可部署的实用方案。

英文摘要

Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is never independently trained at full precision, its bfloat16 checkpoint is a distillation-recovered approximation of the original; QAH distills the 4-bit student directly from the original, uncompressed model. On a GPT-OSS 120B to 60B to MXFP4 pipeline, the QAH student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count, and is released open-weight as Hypernova-60B. Against a matched QAT baseline it reaches a comparable peak about 7 times faster and stays stable under continued training, without hand-tuned early stopping. We also report deployment lessons, including a large, reproducible quality gap between distributed-training backends. Our aim is a recipe deployable without a multi-week hyper-parameter search.

发表机构

  • Multiverse Computing(多元宇宙计算公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑