arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39223cs.LG

QATFactory:面向大规模语言模型量化感知训练与蒸馏的通用、部署对齐框架

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs

Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun

首次发表
浏览论文内容

中文总结 AI 辅助

提出QATFactory框架,通过部署对齐的量化感知蒸馏和强化学习,在BF16训练中模拟低精度量化,支持多种格式,显著提升LLM部署质量,并发现训练策略因格式而异。

中文摘要 AI 辅助

大规模语言模型(LLM)推理日益趋向于采用更低精度以实现硬件加速器的吞吐量,但激进的训练后量化(PTQ)可能会降低模型质量。我们提出了QATFactory,一个开源框架,用于部署对齐的量化感知蒸馏(QAD)和强化学习(QARL)。QATFactory在矩阵乘法以BF16执行的同时模拟部署时的量化,使模型能够适应量化噪声,而无需原生支持目标格式的训练硬件;例如,它支持在缺乏FP4张量核心的H100 GPU上进行NVFP4训练。该框架支持NVFP4、MXFP4以及llama.cpp的Q4_K格式;支持稠密模型和混合专家模型;并支持全参数训练和基于LoRA的训练。它直接将检查点导出到vLLM和llama.cpp,无需额外的有损转换步骤或增加推理开销。利用QATFactory,我们对从8B到230B参数的模型进行了广泛实验,并在生产推理引擎中评估了导出的检查点。在各种模型和格式中,QAD持续提升了部署模型的质量,优于强PTQ基线。在Qwen3.5-9B上,QAD在NVFP4下平均基准准确率达到68.9%,在MXFP4下达到66.0%,分别优于最佳PTQ结果的65.4%和56.4%。通过实验,我们发现尽管两种FP4格式在部署时都量化权重和激活,但最佳训练策略取决于格式:NVFP4通常在训练期间仅量化权重时表现更好,而MXFP4则受益于同时量化权重和激活。在固定的训练token预算下,使用较少的32K序列训练比使用更多的4K序列训练平均准确率提高1.9个百分点。我们发布了完整的QATFactory训练代码和由此产生的检查点。

英文摘要

Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) and reinforcement learning (QARL). QATFactory simulates deployment-time quantization while performing matrix multiplications in BF16, allowing models to adapt to quantization noise without requiring training hardware that natively supports the target format; for example, it supports NVFP4 training on H100 GPUs, which lack FP4 Tensor Cores. The framework supports NVFP4, MXFP4, and $\text{llama}.\text{cpp}$'s Q4_K format; dense and mixture-of-experts models; and both full-parameter and LoRA-based training. It exports checkpoints directly to vLLM and $\text{llama}.\text{cpp}$ without an additional lossy conversion step or added inference overhead. With QATFactory, we conduct extensive experiments on models ranging from 8B to 230B parameters and evaluate exported checkpoints in production inference engines. Across models and formats, QAD consistently improves deployed-model quality over strong PTQ baselines. On Qwen3.5-9B, QAD achieves average benchmark accuracies of 68.9% under NVFP4 and 66.0% under MXFP4, outperforming the best PTQ results of 65.4% and 56.4%, respectively. Through our experiments, we found that although both FP4 formats quantize weights and activations at deployment, the best training strategy is format-dependent: NVFP4 generally performs better when only weights are quantized during training, whereas MXFP4 benefits from quantizing both weights and activations. At a fixed training token budget, training on fewer 32K sequences improves average accuracy by 1.9 points over training on more 4K sequences.

发表机构

  • Together AI
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑