H-Scale:面向NVFP4亚字节大语言模型推理的Hessian引导型尺度优化
H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference
浏览论文内容
中文总结 AI 辅助
针对NVFP4亚字节LLM推理中每组缩放因子优化不足的问题,提出H-Scale方法,利用二阶代理选择硬件合法尺度,改进NVFP4基准并接近BF16参考值。
中文摘要 AI 辅助
NVIDIA Blackwell架构原生支持超精细NVFP4格式,为加速大语言模型(LLM)推理开辟了新机遇。NVFP4的微块设计(如分组大小为16)具备强大的表示灵活性,可捕获局部权重分布并隔离异常值,但也引入了庞大且高度敏感的每组缩放因子空间。现有后训练量化(PTQ)方法主要聚焦于优化量化权重值,却未充分探索该尺度选择步骤。为解决此缺口,我们提出H-Scale,这是一种用于NVFP4每组尺度优化的轻量级后处理方法。H-Scale不最小化普通权重重构误差,而是利用校准激活函数导出的二阶对角代理选择硬件合法的分组尺度,从而更直接地针对层输出扰动。它被设计为多种NVFP4流水线中RTN式尺度选择的即插即用替代方案,仅需适度的离线校准,且在推理时引入严格零开销。在固定评估协议下,对主流LLM的实验表明,H-Scale总体上改进了广泛的NVFP4基准,并使多个变体更接近BF16参考值。
英文摘要
The NVIDIA Blackwell architecture, with native support for the ultra-fine-grained NVFP4 format, opens new opportunities for accelerating large language model (LLM) inference. NVFP4's micro-block design, such as a group size of 16, offers strong representational flexibility for capturing local weight distributions and isolating outliers, but it also introduces a large and highly sensitive space of per-group scaling factors. Existing post-training quantization (PTQ) methods primarily focus on refining quantized weight values, leaving this scale-selection step underexplored. To address this gap, we propose \textbf{H-Scale}, a lightweight post-processing method for NVFP4 per-group scale refinement. Instead of minimizing plain weight reconstruction error, H-Scale selects hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, thereby targeting layer output perturbation more directly. It is designed as a drop-in replacement for RTN-style scale selection in diverse NVFP4 pipelines, requires only modest offline calibration, and introduces strictly zero overhead at inference time. Under a fixed evaluation protocol, experiments on mainstream LLMs show that H-Scale generally improves a broad range of NVFP4 baselines and brings several variants closer to the BF16 reference.
发表机构
- Qwen Team, Alibaba Inc.(通义千问团队,阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。