发表机构
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Zhongguancun Academy; City University of Hong Kong; NJUST(中国科学院自动化研究所; 中国科学院大学人工智能学院; 中关村学院; 香港城市大学; 南京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低比特量化导致的长序列推理退化问题,提出在线策略蒸馏(OPD)阶段,在量化模型实际生成路径上提供教师监督,显著提升数学与代码推理性能保持率。
AI 中文摘要
量化感知蒸馏(QAD)在很大程度上恢复了因低于3比特量化而损失的短问答性能,但仍使数学和代码推理能力大幅受损。长序列生成常常退化为重复循环,在未完成解答的情况下耗尽解码预算。我们将这一差距追溯到量化放大的暴露偏差:QAD在固定语料前缀上训练,而量化引起的偏差会沿模型自身的自回归轨迹累积。为解决这一不匹配,我们引入一个在线策略蒸馏(OPD)阶段,将教师监督置于量化模型实际所到之处。从QAD检查点出发,学生通过部署时使用的量化前向路径生成,并从冻结的全精度教师那里获得关于其自身前缀的反馈,结合密集的令牌级指导与任务验证器奖励。在四个模型上,在2.79和1.88有效比特下,OPD将MATH-500上的平均BF16性能保持率从35%提升至70%,在HumanEval上从66%提升至91%,同时保持短形式性能,且推理增益在匹配预算比较中显著超过持续教师强制的QAD。通过将QAD的稳定低比特初始化与OPD的在线策略推理恢复相结合,我们的框架提供了一个全面的低于3比特解决方案,在恢复长形式推理的同时保持广泛能力。
英文摘要
Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without completing a solution. We trace this gap to quantization-amplified exposure bias: QAD trains on fixed corpus prefixes, while quantization-induced deviations compound along the model's own autoregressive trajectories. To address this mismatch, we introduce an on-policy distillation (OPD) stage that places teacher supervision where the quantized model actually goes. Starting from a QAD checkpoint, the student generates through the quantized forward path used at deployment and receives feedback from a frozen full-precision teacher on its own prefixes, combining dense token-level guidance with task-verifier rewards. Across four models at 2.79 and 1.88 effective bits, OPD raises average BF16 performance retention from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval while preserving short-form performance, with reasoning gains substantially exceeding those of continued teacher-forced QAD in matched-budget comparisons. By coupling QAD's stable low-bit initialization with OPD's on-policy reasoning recovery, our framework provides a comprehensive sub-3-bit solution that preserves broad capabilities while restoring long-form reasoning.
Comments18 pages, 6 figures