arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OSFP4:NVFP4量化的对角平滑与块缩放联合优化

OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization

Neriah Ben David, Ori Meir, Or Ordentlich

arXiv 2610.08231首次发表:更新:

发表机构

School of Computer Science and Engineering; Hebrew University of Jerusalem(计算机科学与工程学院; 耶路撒冷希伯来大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对NVFP4量化精度损失问题,提出OSFP4方案,通过联合优化对角平滑矩阵与块缩放最小化量化误差,在保持高精度的同时保留94-97%的预填充吞吐量。

AI 中文摘要

NVFP4是一种用于大型语言模型(LLM)推理的有吸引力的数据类型,提供紧凑的存储和原生的张量核心加速。然而,使用NVFP4保持精度需要仔细的量化。在这项工作中,我们开发了一种称为NVFP4优化平滑与缩放(OSFP4)的新型量化方案。对于每个线性投影,它使用一个对角平滑矩阵,其条目被优化以最小化NVFP4下的平方矩阵乘积量化误差,同时考虑所使用的舍入过程(要么是舍入到最近,要么是GPTQ风格的连续干扰消除)。这需要对平滑条目以及块缩放进行联合优化,这通过分析乘法抖动FP4量化器而不是固定的确定性量化器来促进。实验表明,OSFP4在相应的量化设置中,在评估的竞争对手中实现了最高的平均精度,同时在测量的工作负载上保留了大约94-97%的供应商NVFP4预填充吞吐量。我们的代码可在以下https URL中获取。

英文摘要

NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). For each linear projection it uses a diagonal smoothing matrix whose entries are optimized to minimize the squared matrix-product quantization error under NVFP4, taking into account the rounding procedure that is used (either round-to-nearest, or GPTQ-style successive interference cancellation). This requires performing joint optimization on the smoothing entries as well as the block scales, which is facilitated by analyzing a multiplicative-dither FP4 quantizer instead of the fixed deterministic one. Experiments show that OSFP4 achieves the highest average accuracy among the evaluated competitors in the corresponding quantization settings, while retaining approximately 94-97\% of vendor NVFP4 prefill throughput on the measured workloads. Our code is available in https://github.com/neriahbd/OSFP4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑