arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReRound:用于解决无校准LLM量化中中点歧义的重构舍入法

ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

He-Yen Hsieh, H. T. Kung

arXiv 2608.11045首次发表:更新:

AI 中文总结

ReRound是一种后训练量化方法,通过条件扩散模型解决无校准LLM量化的中点歧义,在小型LLM的3、4比特量化上优于标准RTN,性能接近依赖校准方法且无额外推理开销。

AI 中文摘要

ReRound(重构舍入)是一种后训练量化方法,用于解决标准就近舍入(RTN)方案在量化量化区间中心附近的权重时固有的中点歧义问题。从预训练的大语言模型(LLM)出发,ReRound训练一个条件扩散模型,以生成该LLM低比特权重的连续重构值。这些重构值作为引导信号,用于消除位于区间中点附近的权重的舍入方向歧义。为了将这种重构引导的舍入与传统RTN相结合,ReRound引入了一个容差度量,用于测量量化权重(而非最终的量化整数)与中点之间的距离:位于中点周围容差区域内的量化权重使用基于扩散的重构值进行量化,而更靠近量化边界的权重则使用RTN进行量化。通过调整容差参数,ReRound生成多个候选量化整数权重矩阵,并选择其反量化权重矩阵候选的主导奇异值与原始全精度权重的主导奇异值最匹配的那个。所选候选确定了ReRound使用的容差参数。ReRound对于较小的LLM特别有效,在一系列此类模型中,它在3比特和4比特权重量化上始终优于标准RTN。ReRound与大量无校准方法相比实现了更优的准确率,与依赖校准的方法相比仍具有竞争力,且完全在离线运行,在低比特推理期间不会引入额外开销。ReRound策略代表了一种低比特量化的新方法,该方法适用于LLM之外的AI模型,本文重点关注其在小型LLM上的应用。

英文摘要

ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs.

Comments16 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑