arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化误差是谱平坦的:单个随机探针是一种校准的、无数据的灵敏度估计器,并应用于预算目标混合精度量化

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

I Kennedy, T Kennedy

arXiv 2609.33923首次发表:更新:

发表机构

Auckland, New Zealand(奥克兰,新西兰)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RAM方法,利用单个随机高斯探针估计量化误差,实现无校准数据的预算目标混合精度量化,在多个模型上优于均匀4位基线。

AI 中文摘要

单个随机高斯探针能够无偏估计层量化误差的Frobenius范数的平方。该估计器表现良好,因为四舍五入到最近值的误差在谱上是平坦的。在来自35B MoE模型和9B稠密模型的1,683个张量中,有效维度是相同形状的i.i.d.噪声值的0.93至0.96倍,并且在MoE上,中位数从2位到8位保持不变。探针变异系数可从张量形状预测。一个探针测量每个张量的灵敏度,误差在4%至7%以内;二十个探针可达到1.3%至1.4%。RAM将这种估计器的传播形式应用于预算目标混合精度量化,无需校准数据。携带网络自身输入统计的高斯探针以六种位宽对每个张量进行评分。背包求解器在精确字节预算下分配位数,并针对灾难性的2位分配设置防护措施。一次探针通过即可满足任何预算。孤立和传播的分数在Qwen3.5-35B-A3B上独立地对张量进行排序(Spearman相关系数为-0.01),然而传播的探针与来自真实激活的GPTQ层目标排名相关0.81至0.83,而孤立估计器与其不相关。该目标是错误的分配目标:在Qwen3.8-27B上匹配字节时,块输出探针优于供应商的IQ3_M混合和从真实激活分配的oracle。在Qwen3-8B上,传播的探针在匹配字节时与HAWQ-V2持平。在从8B到122B的七个架构中,探针计时在单台工作站上九分钟内完成400B模型,RAM在测试的MoE模型上比大小相当的均匀4位构建实现了3.5%至13.6%更低的WikiText-2困惑度中位数。

英文摘要

A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96 times the i.i.d. noise value of the same shape, and on the MoE the median is unchanged from 2-bit to 8-bit. The probe coefficient of variation is predictable from tensor shape. One probe measures per-tensor sensitivity to within 4 to 7%; twenty probes reach 1.3 to 1.4%.RAM applies the propagated form of this estimator to budget-targeted mixed-precision quantization with no calibration data. Gaussian probes carrying the network's own input statistics score every tensor at six bit-widths. A knapsack solver allocates bits under an exact byte budget, with guardrails against catastrophic 2-bit assignments. One probe pass serves any budget. Isolated and propagated scores rank tensors independently on Qwen3.5-35B-A3B (Spearman -0.01), yet the propagated probe rank-correlates 0.81 to 0.83 with the GPTQ layer objective from real activations, while the isolated estimator is uncorrelated with it. That objective is the wrong allocation target: at matched bytes on Qwen3.8-27B, a block-output probe beats a vendor IQ3_M mix and an oracle that allocates from the real-activation objective. On Qwen3-8B the propagated probe ties HAWQ-V2 at matched bytes. Across seven architectures from 8B to 122B, with probe timing up to a 400B model in nine minutes on one workstation, RAM reaches 3.5 to 13.6% lower median WikiText-2 perplexity than size-comparable uniform 4-bit builds on the tested MoE models. (Black Sheep Ai baa.ai)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑