发表机构
City University of Hong Kong; The University of Hong Kong(香港城市大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将权重平均建模为训练进度与鲁棒性的权衡,推导出统一平均核族,为训练后量化下的权重平均提供理论指导,并实验验证其有效性。
AI 中文摘要
大语言模型(LLMs)通常以高精度进行预训练,但越来越多地以低精度的训练后量化(PTQ)进行部署。近期研究表明,在预训练期间使用权重平均可以比学习率衰减更好地提升PTQ性能,这表明它可能为改善从预训练到量化的过渡提供一种简单方法。然而,权重平均背后的机制尚未得到充分解释,这导致性能提升不一致且脆弱,从而阻碍了从业者自信地应用此类技术。为此,我们将权重平均表述为在保留训练进度与提升扰动鲁棒性之间的权衡。我们进一步推导出一个连续的平均核族,该核族统一了传统策略,并在两个竞争目标之间实现了帕累托前沿。关键的是,我们开发了一个在PTQ下执行权重平均的理论框架。可以证明,较粗的量化对扰动更敏感,而较细的量化可能受影响较小。因此,我们的结果可以为在不同PTQ条件下执行权重平均提供统一的理论指导。实验验证了预测行为以及所提出的平均策略。代码可在以下网址获取:此https URL。
英文摘要
Large language models (LLMs) are typically pretrained in high precision but increasingly deployed with low-precision post-training quantization (PTQ). Recent studies have shown that using weight averaging during pretraining can improve PTQ performance compared with learning-rate decay, suggesting that it might provide a simple way to improve the pretraining-to-quantization transition. But the mechanism behind weight averaging remains insufficiently explained. This leads to inconsistent and fragile performance gains, thereby preventing practitioners from applying such a technique confidently. As a response, we formulate weight averaging as a trade-off between retaining training progress and improving robustness under perturbation. We further derive a continuous family of averaging kernels that unifies conventional strategies and achieves the Pareto frontier between the two competing goals. Critically, a theoretical framework for performing weight averaging under PTQ is developed. It can be shown that coarser quantization is more susceptible to perturbations, whereas finer quantization could be less affected. Thus, our results could provide unified theoretical guidance for performing weight averaging under different PTQ conditions. Experiments validate both the predicted behavior and the proposed averaging strategy. Code is available at https://github.com/MOFA-LAB/weight-averaging-for-ptq.