arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Softmax 重参数化用于输出头量化

Softmax Reparameterization for Output-Head Quantization

Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King

arXiv 2609.31291首次发表:更新:

发表机构

Adobe SDC(Adobe SDC)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对小语言模型输出头量化导致的性能下降,提出 softmax 重参数化方法,通过一维搜索选择等效输出头,在 W4 量化下显著降低 KL 散度,且不增加推理开销。

AI 中文摘要

大词汇表使得输出头在小语言模型中成为可观的推理成本。我们提出了 softmax 重参数化,一种训练后方法,在量化之前选择一个功能等效的输出头。该方法从每个输出行中减去词汇表行均值的标量倍数,并通过验证 KL 分别针对 RTN、激活加权 MSE 和全 Hessian GPTQ 选择系数。这个一维搜索包括原始头和固定均值中心化,保留全精度 softmax 分布,并保持训练后的解码器不变;秩一校正处理非线性 logit 路径,如 soft-capping。在七个头上,W4 的收益集中在基线量化显著扭曲预测的地方:在 Phi-4-mini 上,AW-MSE KL 从 0.936 降至 0.256。这些收益在更强的 GPTQ 校准下依然存在,并且与精确的逐通道缩放和仿射量化互补。在四个头和三个 W4 量化器上,冻结的 WikiText 选择的系数也迁移到 C4 和 OpenWebMath,在冻结系数与 1 不同的所有 18 个比较中优于均值中心化,并在其余六个比较中与之匹配。在 W2 下,作为压缩压力测试,收益在几乎整个模型-量化器矩阵上扩大。匹配残差分析表明,改进的保真度可以伴随更大的 logit 重建误差,同时降低残差的 Fisher 加权成本。对于移位兼容的头,重参数化不增加推理操作,并保持打包的 W4 执行:在解码器保持 BF16 的情况下,量化 Phi 输出头相对于 BF16 头基线将批量一代延迟降低了 10.8%。

英文摘要

Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL. For linear-softmax heads, these shifts preserve full-precision predictions exactly and require no decoder retraining; a rank-one correction extends the construction to nonlinear logit paths. Across seven output heads and three quantizers, W4 gains are largest where baseline quantization substantially distorts predictions: test KL falls by 93% on XGLM under RTN and by 73--77% on Phi, BLOOM, and BLOOMZ under activation-weighted MSE. Heads with low baseline error change little; at W2, used as a compression stress test, benefits extend more broadly. On Phi, the gains persist under stronger GPTQ calibration; a separate untouched holdout reproduces the improvements on Phi and BLOOM. Frozen WikiText-selected coefficients also transfer without retuning to C4 and OpenWebMath. Residual analysis on Phi shows how fidelity can improve despite greater total logit error: the selected representative reduces error on likely outputs and lowers its Fisher-weighted cost. For shift-compatible heads, the shift adds no inference operation. With the decoder held in BF16, a packed W4 Phi output head reduces batch-one generation latency by 10.8%, and reparameterization preserves this speedup.

Comments32 pages, including appendix, ICLR 2027 submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑