OnlineQAT:面向超低位大型语言模型的在线策略蒸馏
OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
浏览论文内容
中文总结 AI 辅助
OnlineQAT通过两阶段框架,先分块QAT初始化,再对学生生成响应进行在线策略蒸馏,利用冻结全精度教师的反向KL信号,在Qwen3-1.7B上取得W3A16和W2A16最佳平均性能,优于固定补全训练。
中文摘要 AI 辅助
量化感知训练(QAT)可以在大型语言模型被压缩到四比特以下时恢复大部分精度损失。然而,现有的恢复阶段通常是在固定补全或教师生成的答案上进行优化的,而部署的量化模型则基于其自身生成的前缀进行条件化。因此,量化误差可能将模型推入离线恢复数据中不存在的状态。我们提出了OnlineQAT,这是一个两阶段框架,首先通过分块QAT获得可用的低位初始化,然后在学生生成的响应上执行在线策略蒸馏(OPD)。在每个访问到的前缀处,冻结的全精度教师提供采样的反向KL训练信号。在Qwen3-1.7B上,OnlineQAT在比较的量化方法中取得了最佳平均值:W3A16下为57.28,W2A16下为32.52,分别比ReasoningQAT提高了2.90和0.44个点。结果表明,学生访问的状态提供了超越固定补全训练的有用恢复信号,尤其是在三比特设置下。
英文摘要
Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Sun Yat-sen University(中山大学)
- PolyU-Daya Bay Technology and Innovation Research Institute(香港理工大学大亚湾技术创新研究院)
机构由 AI 辅助整理,请以论文原文为准。