arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22106cs.LGcs.CV

PRQuant:用于低开销推理的排列残差量化

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

  • Huawei Technologies Co., Ltd.(华为技术有限公司)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Yuantian Shao, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang

AI总结:

PRQuant通过通道排列与静态残差补偿,实现无训练低开销的低比特量化推理,在多个基准上优于MXFP4。

AI中文摘要:

线性层低比特量化的精度往往受少数离群值主导。尽管现有方法(如平滑、旋转或基于残差的方法)可能缓解此问题,但它们常常给权重引入新的精度瓶颈。此外,这些技术大多以在线方式实现,可能导致沉重的执行开销。为解决上述问题,我们提出PRQuant(排列残差量化),一个无需训练且低开销的框架,结合了通道重组与静态权重侧残差补偿。在AWQ风格缩放后,PRQuant识别对权重量化误差贡献最大的输入通道,将它们排列到连续的尾部块中,并离线构建其残差权重子张量。在推理期间,这种连续结构使激活侧无需昂贵的在线收集操作即可无缝使用尾部块,并将分散的残差补偿转化为规则的尾部增广GEMM,大幅降低延迟。实验表明,PRQuant有效降低下投影重建误差。消融研究证实,平滑和残差补偿是数值改进的主要驱动力,而排列提供一致的边际数值收益,更重要的是,它实现了硬件友好的连续布局,消除了动态收集开销。总体而言,PRQuant在五个下游基准的平均准确率上优于默认MXFP4和所评估的PTQ基线,在Qwen3-4B-Instruct-2507和Qwen3-30B-A3B-Instruct-2507上分别比MXFP4提高1.24和0.55。

英文摘要:

Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (Permutation Residual Quantization), a training-free framework that combines channel permutation with offline weight residual compensation. PRQuant identifies the scaled-weight columns with the largest quantization errors and permutes them into contiguous tail blocks. This structure allows the corresponding weight residuals to be precomputed entirely offline, while replacing scattered activation gathering with simple contiguous access during inference, yielding a single regular MXFP4 GEMM for compensated computation. Experiments show that PRQuant substantially reduces down-projection reconstruction error, with scaling and residual compensation providing the main numerical gains while permutation enables a hardware-friendly contiguous layout. Comprehensive experiment results on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 illustrate that PRQuant achieves up to averagelly 2.6x and 1.8x operator speedup over BF16 respectively, while preserving near plain MXFP4 end-to-end decoding efficiency. Across five downstream benchmarks, PRQuant achieves the best average accuracy among the quantized methods, improving accuracy over MXFP4 by 1.24 and 0.55, respectively.

↑