发表机构
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型量化中预算变化及层敏感度评估问题,提出MixQuant自适应框架,通过边缘化失真获与预算无关分数、校准参数并惩罚不合理分配,在多种模型上优于基线,提升准确率并降低困惑度。
AI 中文摘要
混合精度量化通过为敏感层分配更高的比特宽度来提高训练后量化的准确性,但现有方法是针对单个固定内存预算解决分配问题。实际中预算在不同部署中变化且在校准时尚未知。自适应量化通过一次离线校准解决此问题,但当前方法评估层敏感度时未考虑其对其他层量化级别的依赖性。我们表明层的敏感度强烈依赖于其上游层的比特宽度,且这种依赖会改变最终的首选比特分配。我们提出MixQuant,一个与技术无关的自适应框架,它可以包装任何基础量化器。MixQuant通过对随机量化的上游配置边缘化每层的失真来获得与预算无关的分数,在校准器自身生成的计划上校准量化器参数,并惩罚使层处于最低比特宽度的分配。单次贪心遍历即可在部署时适应任何预算。在AWQ和GPTQ下的Llama-3.2-3B、Llama-2-7B和Mistral-7B模型中,MixQuant在每种设置下均优于自适应和混合精度基线,在最紧预算下平均准确率提高多达8个点,困惑度从12.43降至10.70,同时以可忽略的部署成本匹配整数线性规划求解器。
英文摘要
Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget. In practice the budget varies across deployments and is unknown at calibration time. Adaptive quantization addresses this with one offline calibration that serves any budget, yet current methods score layer sensitivity in a manner that does not consider its dependency on quantization levels of other layers. We show that a layer's sensitivity depends strongly on the bitwidths of its upstream layers and that this dependence shifts the resulting preferred bit allocation. We propose MixQuant, a technique-agnostic adaptive framework that wraps any base quantizer. MixQuant marginalizes each layer's distortion over random quantized upstream configurations to obtain budget-agnostic scores, calibrates the quantizer's parameters on plans the allocator itself produces, and penalizes allocations that leave layers at the lowest bitwidths. A single greedy pass then serves any budget at deployment. Across Llama-3.2-3B, Llama-2-7B, and Mistral-7B under AWQ and GPTQ, MixQuant outperforms adaptive and mixed-precision baselines in every setting, improving average accuracy by up to 8 points and reducing perplexity from 12.43 to 10.70 at the tightest budget, while matching an ILP solver at negligible deployment cost.
CommentsPreprint