arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.09555cs.CLcs.AI

在微缩浮点格式下对大语言模型进行后训练量化评估

Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats

  • Huawei Technologies(华为技术)

机构由 AI 辅助整理,请以论文原文为准。

Manyi Zhang, Ji-Fu Li, Zhongao Sun, Haoli Bai, Hui-Ling Zhen, Zhenhua Dong, Xianzhi Yu

更新

AI总结:

本文研究了在微缩浮点格式下大语言模型后训练量化的效果,发现MXFP8性能接近无损,而MXFP4存在显著精度损失,且格式兼容性对PTQ效果影响显著。

AI中文摘要:

微缩浮点(MXFP)作为一种低精度格式,已经 emerged 为大语言模型(LLMs)的有前景的格式。尽管已经提出了各种后训练量化(PTQ)算法,但它们大多集中在整数量化上,而它们在MXFP格式下的适用性和行为仍然鲜为人知。为了解决这一差距,本文系统地研究了在MXFP格式下的PTQ,涵盖了超过7种PTQ算法、15个评估基准和3个LLM家族。关键发现包括:1)MXFP8始终实现接近无损性能,而MXFP4引入了显著的准确性退化并仍具挑战性;2)在MXFP下PTQ的有效性强烈依赖于格式兼容性,某些算法范式比其他范式更有效;3)PTQ性能在模型家族和模态之间表现出高度一致的趋势,特别是,在多模态LLMs中,量化敏感性主要由语言模型而非视觉编码器主导;4)量化缩放因子是MXFP4中的关键误差源,简单的预缩放优化策略可以显著减轻其影响。这些结果为适应现有PTQ方法到MXFP量化提供了实用指导。

英文摘要:

Microscaling Floating-Point (MXFP) has emerged as a promising low-precision format for large language models (LLMs). Despite various post-training quantization (PTQ) algorithms being proposed, they mostly focus on integer quantization, while their applicability and behavior under MXFP formats remain largely unexplored. To address this gap, this work conducts a systematic investigation of PTQ under MXFP formats, encompassing over 7 PTQ algorithms, 15 evaluation benchmarks, and 3 LLM families. The key findings include: 1) MXFP8 consistently achieves near-lossless performance, while MXFP4 introduces substantial accuracy degradation and remains challenging; 2) PTQ effectiveness under MXFP depends strongly on format compatibility, with some algorithmic paradigms being consistently more effective than others; 3) PTQ performance exhibits highly consistent trends across model families and modalities, in particular, quantization sensitivity is dominated by the language model rather than the vision encoder in multimodal LLMs; 4) The scaling factor of quantization is a critical error source in MXFP4, and a simple pre-scale optimization strategy can significantly mitigate its impact. Together, these results provide practical guidance on adapting existing PTQ methods to MXFP quantization.

↑