EfficientQAT:面向大语言模型的高效量化感知训练
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
- The University of Hong Kong(香港大学)
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大语言模型量化感知训练资源消耗过高的问题,提出包含全参数分块训练与量化参数端到端训练两阶段的EfficientQAT算法,在多种规模与类型的LLM低比特量化任务中性能优于现有方法,且训练资源需求大幅降低。
AI中文摘要:
大语言模型(LLM)是现代自然语言处理与人工智能领域的核心技术,但其内存占用过高的问题带来了部署挑战。量化感知训练(QAT)通过低比特表示降低内存消耗且精度损失极小,是可行的解决方案,但所需训练资源过多,实际应用难度大。针对这一问题,我们提出高效量化感知训练算法EfficientQAT,大幅提升了QAT的落地可行性。EfficientQAT包含两个连续阶段:全参数分块训练(Block-AP)与量化参数端到端训练(E2E-QP)。据我们所知,Block-AP是首个实现全参数分块直接训练的方法,通过拓展优化过程的解空间,降低了低比特场景下的精度损失。随后E2E-QP仅对量化参数(步长)进行端到端训练,通过考量所有子模块间的交互进一步提升量化模型性能。大量实验表明,在参数规模7B到70B的基础LLM、指令微调LLM、多模态LLM等多种模型、不同量化比特设置下,EfficientQAT的表现均优于现有量化方法。例如,EfficientQAT仅用单张A100-80GB GPU、耗时41小时即可训练出2比特Llama-2-70B模型,与全精度模型相比精度下降不到3个点(69.48 vs. 72.41)。代码已开源至https://github.com/OpenGVLab/EfficientQAT。
英文摘要:
Large language models (LLMs) are crucial in modern natural language processing and artificial intelligence. However, they face challenges in managing their significant memory requirements. Although quantization-aware training (QAT) offers a solution by reducing memory consumption through low-bit representations with minimal accuracy loss, it is impractical due to substantial training resources. To address this, we propose Efficient Quantization-Aware Training (EfficientQAT), a more feasible QAT algorithm. EfficientQAT involves two consecutive phases: Block-wise training of all parameters (Block-AP) and end-to-end training of quantization parameters (E2E-QP). To the best of our knowledge, Block-AP is the first method to enable direct training of all parameters in a block-wise manner, reducing accuracy loss in low-bit scenarios by enhancing the solution space during optimization. E2E-QP then trains only the quantization parameters (step sizes) end-to-end, further improving the performance of quantized models by considering interactions among all sub-modules. Extensive experiments demonstrate that EfficientQAT outperforms previous quantization methods across a range of models, including base LLMs, instruction-tuned LLMs, and multimodal LLMs, with scales from 7B to 70B parameters at various quantization bits. For instance, EfficientQAT obtains a 2-bit Llama-2-70B model on a single A100-80GB GPU in 41 hours, with less than 3 points accuracy degradation compared to the full precision (69.48 vs. 72.41). Code is available at https://github.com/OpenGVLab/EfficientQAT.