SoftWater:面向Softmax量化的类别感知速率分配
SoftWater: Class-Aware Rate Allocation for Softmax Quantization
浏览论文内容
中文总结 AI 辅助
针对小型LLM的Softmax层量化问题,提出类别感知速率分配方法SoftWater,在多模型上优于现有量化器,可大幅降低输出层KL散度,使输出层量化的困惑度损失可控。
中文摘要 AI 辅助
训练后量化流程通常将Softmax输出层保留为高精度。然而在带有现代词表的小型大语言模型(LLM)中,输出层(head)占所有参数的15%–30%,因此标称的“2比特”模型若使用fp16精度的输出层,其每个权重可存储的比特数会是前者的数倍。我们将Softmax层量化建模为原始与量化输出分布间KL散度约束下的速率-失真问题。二阶分析揭示了类别感知几何:量化误差由特征协方差与类别特定的Softmax曲率共同加权。可分性近似将Kn×Kn的Cholesky分解替换为按类别缩放的n×n分解,使格点可通过连续干扰消除(SIC)编码,且所有统计量来自单次前向传播。所得方法SoftWater为高频低方差类别提供精细网格,为稀有类别提供粗糙网格,在Zipfian token分布下存在显著差距。在1B至32B的5个模型上,SoftWater在60个测试点中的59个上,于匹配的输出层速率下,优于已发布的WaterSIC量化器(其在线性层WMSE下接近最优,但在输出层KL散度下非最优),且未使用该流程的任何改进措施,在2比特设置下将输出层诱导的KL散度降低了6.5倍至8.3倍。对于量化主体的Llama-3.2-1B-Instruct,2比特输出层可消除45%–60%的存储字节,仅带来2.9%–3.7%的困惑度(perplexity)提升。由于类别侧统计量来自校准数据,将校准数据与部署域匹配可在该域始终获得最低KL散度。在绑定(tied)模型上,4比特输出层接近无损,2比特输出层的困惑度损失低于4%,使此类模型的输出层量化具备可行性。
英文摘要
Post-training quantization pipelines routinely leave the softmax output layer in high precision. Yet in small LLMs with modern vocabularies, the head holds 15--30\% of all parameters, so a nominal ``2-bit'' model with an fp16 head can store several times as many bits per weight. We pose softmax-layer quantization as a rate-distortion problem under the KL divergence between the original and quantized output distributions. A second-order analysis reveals a class-aware geometry: quantization error is weighted jointly by feature covariance and class-specific softmax curvature. A separability approximation replaces the $Kn\times Kn$ Cholesky with one $n\times n$ factorization rescaled per class, making the lattice encodable by successive interference cancellation, with both statistics from a single forward pass. The resulting method, SoftWater, gives fine grids to frequent, low-variance classes and coarse grids to rare ones, a large gap under Zipfian token distributions. Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer (near-optimal under linear-layer WMSE but not output KL) at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times$ at 2 bits. On Llama-3.2-1B-Instruct with quantized bodies, a 2-bit head removes 45--60\% of stored bytes for a $2.9$--$3.7\%$ perplexity increase. Because the class-side statistic comes from calibration data, matching calibration to the deployment domain gives the lowest KL on that domain throughout. On a tied model, a 4-bit head is near-lossless and a 2-bit head costs under 4\% perplexity, making head quantization of such models practical.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。