发表机构
McGill University(麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
JARQ是一种即插即用的量化细化方法,通过联合交替优化组尺度和编码,提升大语言模型量化后的性能,降低困惑度并提高准确率。
AI 中文摘要
针对大型语言模型的组后训练量化器将权重舍入到一个网格上,该网格不会重新拟合到由此产生的整数编码。我们表明,这留下了准确性:最佳网格取决于编码,输入相关性耦合了不同组的误差,有用的编码更改通常同时涉及许多编码。我们提出了JARQ,一种即插即用的细化方法,从任何组量化器开始,交替进行所有组尺度的联合最小二乘拟合与有界Babai提议,这些提议在当前网格上同时移动组的许多编码。该问题是一个双线性盒约束混合整数最小二乘问题;求解器无需反向传播,在精确尺度求解下不会增加逐层目标,并保持主机的位宽、组、零点和推理成本。在Llama-2、Llama-3和Qwen模型中,使用RTN、GPTQ、OmniQuant和AWQ作为主机,JARQ在96次比较中降低了90次的困惑度,将三比特RTN困惑度降低了高达36%,在24种配置中提高了23种的平均多项选择准确率,并改善了QEP、QuaRot和OJBKQ的输出,每个7B块的处理时间不到一分钟。
英文摘要
Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve many codes at once. We propose JARQ , a plug-in refinement that starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with bounded Babai proposals that move many codes of a group together on the current grid. The problem is a bilinear box-constrained mixed-integer least-squares problem; the solver is backpropagation-free, does not increase the layer-wise objective under exact scale solves, and keeps the host's bit width, groups, zero points, and inference cost. Across Llama-2, Llama-3, and Qwen models with RTN, GPTQ, OmniQuant, and AWQ hosts, JARQ lowers perplexity in 90 of 96 comparisons, cuts three-bit RTN perplexity by up to 36%, raises mean multiple-choice accuracy in 23 of 24 configurations, and improves QEP, QuaRot, and OJBKQ outputs, at under a minute per 7B block.