FOCUS:通过耦合松弛与双粒度缩放实现FP4优化
FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
FOCUS是一种后训练量化框架,通过耦合松弛与双粒度缩放优化FP4,在不增加推理开销的情况下,于MXFP4和NVFP4格式下实现了LLM的最优FP4精度。
AI中文摘要:
大语言模型(LLMs)性能优异,但因规模庞大部署成本高昂。FP4量化(含MXFP4、NVFP4等格式)是现代加速器原生硬件支持的理想方案,但FP4精度下保持准确性仍具挑战。关键瓶颈在于缩放优化:现有方法将量化与反量化缩放紧密耦合,要求二者符合硬件所需的离散低精度格式(如MXFP4中的E8M0),而量化缩放从未被存储且无需遵守该约束,存在未开发的优化空间。本研究提出FOCUS,一种后训练量化框架,通过耦合松弛与双粒度缩放实现FP4的端到端缩放学习。耦合松弛缩放(CRS)以可学习的全精度系数放松量化与反量化缩放的紧密耦合,在不破坏硬件兼容性的前提下实现更有效的优化;双粒度缩放(DGS)进一步以更精细的子块粒度优化量化缩放,使其更精准适配局部权重分布。在多个LLMs系列与基准上的实验显示,FOCUS在MXFP4和NVFP4格式下均实现了最先进的FP4精度,且无额外推理开销。代码与量化模型将在此URL发布。
英文摘要:
Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers an appealing solution with native hardware support on modern accelerators. However, maintaining accuracy under FP4 precision remains difficult. A key bottleneck lies in scale optimization: existing methods tightly couple the quantization and dequantization scales, forcing both to conform to the discrete low-precision format required by hardware, such as E8M0 in MXFP4. Yet the quantization scale is never stored and need not obey this constraint, suggesting a significant untapped optimization space. In this work, we propose FOCUS, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling. Coupled-Relaxation Scaling (CRS) relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimization without breaking hardware compliance. Dual-Granularity Scaling (DGS) further refines the quantization scale at a finer sub-block granularity, allowing more precise adaptation to local weight distributions. Experiments across multiple LLM families and benchmarks show that FOCUS achieves state-of-the-art FP4 accuracy under both MXFP4 and NVFP4 formats, while introducing no additional inference overhead. Code and quantized models will be released at https://github.com/tencent/AngelSlim.