arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36654cs.LGcs.CLcs.DC

重放曲率:面向大语言模型推理的准确且可扩展的NVFP4量化方法

Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

Ruiyi Ding, Jie Li, Kang He, Ziyan Liu, Chengru Song, Yuedong Xu, Yuan Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型推理中的NVFP4量化,提出Schur Replay算法和分层执行基础设施,实现高精度恢复并显著降低每层耗时和内存占用。

中文摘要 AI 辅助

大语言模型使得权重存储和内存流量成为主要的推理成本,从而推动了以仅用几个比特表示每个权重的低精度格式的发展。此类格式使用缩放因子将浮点值映射到一个小型码本;NVFP4通过让每16个E2M1权重共享一个E4M3块缩放因子来改善局部范围利用率。在GPTQ中选择该缩放因子较为困难,因为量化一列会更新后续列,因此独立评估一个块可能会错误估计其最终重建误差。大型模型带来了第二个挑战:全精度权重、校准激活和二阶状态无法全部驻留在单个加速器上,而将完整层分配给设备会导致每个耗时的层串行求解。我们引入了\emph{Schur Replay},一种缩放因子选择算法,它重现每个块缩放因子引起的GPTQ更新,并在考虑未量化列的补偿后对所得的块误差进行评分。此外,我们的执行基础设施仅保留活动层驻留,将激活分层存储于设备、主机和磁盘,导出后淘汰全精度层,并将独立的输出行分布到张量并行秩中。算法与基础设施共同在Qwen3.5-397B-A17B和Llama-3.3-70B-Instruct上的七个基准中实现了从BF16的99.35%和100.84%的问题加权恢复率。在397B模型上,该基础设施相比ModelOpt将每层实测时间减少了15.17倍,相比LLM Compressor减少了23.14倍,且每GPU内存使用更低。

英文摘要

Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosing that scale is difficult in GPTQ because quantizing one column updates those that follow, so evaluating a block independently can misestimate its final reconstruction error. Large models pose a second challenge: full-precision weights, calibration activations, and second-order state cannot all remain on one accelerator, while assigning complete layers to devices leaves each time-consuming layer solve serial. We introduce \emph{Schur Replay}, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns. Separately, our execution infrastructure keeps only the active layer resident, tiers activations across device, host, and disk, retires full-precision layers after export, and distributes independent output rows across tensor-parallel ranks. Together, the algorithm and infrastructure attain $99.35\%$ and $100.84\%$ question-weighted recovery from BF16 across seven benchmarks on Qwen3.5-397B-A17B and Llama-3.3-70B-Instruct. On the 397B model, the infrastructure reduces measured per-layer time by $15.17\times$ over ModelOpt and $23.14\times$ over LLM Compressor, with lower memory used per GPU.

发表机构

  • KlingAI Research(可灵AI研究院)
  • Fudan University(复旦大学)
  • Shanghai AI Incubation and Innovation Center(上海人工智能孵化与创新中心)
  • Shanghai Academy of AI for Science(上海人工智能科学研究院)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑