用于少比特整数的带符号对称量化
Signed Symmetric Quantization for Few-Bit Integers
- AMD(超威半导体公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究少比特整数量化问题,提出带符号对称量化方法,通过有原则的符号选择规则放置额外可表示值,保持对称量化运行时特征,理论分析其在量化误差上有条件最优,实证验证在模型上有效果提升。
AI中文摘要:
带符号整数字母表中可表示的负值比正值多一个。然而,标准对称整数量化器按惯例将其比例固定为严格正值,这会将这个额外的可表示值分配给负尾,并可能导致正异常值的裁剪。在这项工作中,我们表明,在少比特精度下,这种裁剪是量化误差的一个重要来源。非对称量化通过一个零点解决了这个问题,将网格向观测数据范围移动;然而,这种灵活性会带来运行时的代价。例如,在特定CPU上,4比特对称格式比非对称格式使用的内存少9%,吞吐量高2.45倍。我们强调带符号对称量化是第三种选择,它保留了对称量化的运行时特征,而没有非对称格式的代价:我们的带符号绝对最大值网格通过一个有原则且轻量级的符号选择规则将额外的可表示值放在主导异常值尾上,同时将零点保持在零。我们的理论分析提供了两个主要结果。首先,我们确定带符号绝对最大值网格在$\ell_2$量化误差上有条件地最优,并表明在低比特宽度下,该条件对预训练大语言模型中88 - 99%的权重组成立。其次,我们表明将标准对称量化器的比例取反在分析上等同于在相同带符号整数字母表上的单位零点偏移。我们在Qwen3、Qwen3.5和Llama3系列模型上对我们的提议进行了实证验证,并且在没有额外推理成本的情况下,观察到与标准无符号对称量化器相比,困惑度和下游少样本准确率有所提高。
英文摘要:
The signed integer alphabet contains one more negative representable value than positive. Yet, by convention, the standard symmetric integer quantizer fixes its scale to be strictly positive, which assigns this extra representable value to the negative tail and can force clipping of positive outliers. In this work, we show that, at few-bit precision, such clipping is a non-trivial source of quantization error. Asymmetric quantization addresses this problem with a zero point, shifting the grid toward the observed data range; however, this flexibility is well-known to carry a runtime penalty. For example, in llama.cpp on an AMD EPYC(TM) "Turin" CPU, a 4-bit symmetric format uses up to 9% less memory with up to 2.45$\times$ higher throughput than its asymmetric counterpart. We highlight signed symmetric quantization as a third option that retains the runtime profile of symmetric quantization without the penalty of the asymmetric format: our signed absmax grid places the extra representable value on the dominant-outlier tail through a principled and lightweight sign selection rule while keeping the zero point at zero. Our theoretical analysis offers two main results. First, we establish the signed absmax grid as conditionally bound-optimal on $\ell_2$ quantization error, and show that the condition holds for 88-99% of weight groups across pre-trained large language models (LLMs) at low bit widths. Second, we show that negating the scale of a standard symmetric quantizer is analytically equivalent to a unit zero point shift on the same signed integer alphabet. We empirically validate our proposal on models from the Qwen3, Qwen3.5, and Llama3 families, and observe improvement in perplexity and downstream few-shot accuracy over the standard unsigned symmetric quantizer at no extra inference cost