Q-PACE:量化感知训练中的动态精度分配
Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training
浏览论文内容
中文总结 AI 辅助
针对量化感知训练中激进量化损害性能的问题,提出Q-PACE方法,通过二阶敏感性模型动态分配层精度,在低内存预算下保持损失性能,并在40亿参数LLM上验证其有效性。
中文摘要 AI 辅助
量化感知训练(QAT)利用低精度算术来降低大型语言模型(LLM)部署的成本,但激进的量化会降低最终模型性能。一种常见的补救措施是混合精度训练,即为部分层分配高精度以保持性能,同时控制成本。这种方法需要在训练期间为模型层分配精度。我们提出了一种新方法,称为Q-PACE,它包含一个二阶敏感性模型,该模型将损失增加预测为按层曲率系数加权的量化噪声均方误差(MSE)之和。在训练期间,我们定期使用跨层的扰动重新计算这些系数,并重新分配精度。在参数规模高达40亿的LLM上的预训练和监督微调实验表明,Q-PACE始终优于现有的混合精度训练方案,并在显著更低的总内存预算下实现相当的损失。我们进一步发现,量化敏感性高度可预测,取决于层深度和层类型,并且其在训练期间的稳定性允许进行不频繁、低成本的重新校准。
英文摘要
Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraining and supervised fine-tuning experiments on LLMs of up to 4B parameters show that Q-PACE consistently improves over existing mixed-precision training recipes, and achieves comparable loss at substantially lower total memory budgets. We further find that quantization sensitivity is highly predictable by depth and layer type, and its stability during training allows for infrequent, cheap recalibration.
发表机构
- ISTA
- EPFL
机构由 AI 辅助整理,请以论文原文为准。