arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

饱和使量化误差具有可加性:一种带有证明的覆盖模型

Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate

Joshua Hill

arXiv 2607.12266首次发表:更新:

发表机构

Baseten Labs, Inc.(巴斯特恩实验室公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究混合精度量化中如何确定模型高精度部分,通过分析量化损失变化,发现大部分方差由每层效应解释,提出覆盖模型,该模型支持两个预测器,加法模型是最优一阶预测器,覆盖模型本身也有优势,在特定参数模型上实现低KL散度且能解决相关任务。

AI 中文摘要

混合精度量化必须决定模型的哪些部分保持更高精度。基于灵敏度的方法(如HAWQ和CoopQ)的一个共同前提是,量化一组层的损失可以从单独测量的每层或成对灵敏度重建。我们在当前部署的4位权重和激活精度下测试这个前提,将量化层集\(S\)的损失\(f(S)\)的变化视为布尔立方体上的集函数,并通过两种经典的基变换进行分析。此分析得出两个发现。首先,在从部署分布中抽取的配置中,\(f\)的方差的85%-93%仅由每层效应解释。其次,每层项之和的单调变换再现了\(f\)对配置的排序,最多错排2%的对。我们提出覆盖模型\(f(S)=c(1 - \prod_{i\in S}(1 - a_i))\),它在拟合的\(L\)个断点处将\(f\)的测量方差轮廓再现到几个百分点以内。这种结构支持配置损失的两个预测器,每个预测器有\(L + 1\)个参数。加法模型是最优的一阶预测器。根据帕塞瓦尔恒等式,其均方误差等于未由每层效应解释的\(f\)的方差,我们在全格上测量,在全网络规模上进行样本外估计,并在每个结果中报告,作为任何加法模型性能的证明。覆盖模型本身是第二个预测器。作为匹配内存的分配器,它们在30B到355B参数的模型上,在所比较的分配器中实现了最低的KL散度。低于4位时,在梯度灵敏度的分配不再产生终止世代的预算下,所得分配继续解决编码和推理任务。

英文摘要

Mixed-precision quantization must decide which parts of a model to keep at higher precision. A common premise, shared by sensitivity-based methods such as HAWQ and CoopQ, is that the loss from quantizing a set of layers can be reconstructed from per-layer or pairwise sensitivities measured in isolation. We test this premise at the 4-bit weight-and-activation precisions now being deployed, treating the change in loss $f(S)$ from quantizing a layer set $S$ as a set function on the Boolean cube and analyzing it through two classical changes of basis. This analysis yields two findings. First, across configurations drawn from the deployment distribution, 85--93\% of the variance of $f$ is explained by per-layer effects alone. Second, a monotone transform of a sum of per-layer terms reproduces $f$'s ranking of configurations, misordering at most 2\% of pairs. We propose the coverage model $f(S)=c\bigl(1-\prod_{i\in S}(1-a_i)\bigr)$, which reproduces the measured variance profile of $f$ to within a few percent from its $L$ fitted break-rates. This structure supports two predictors of a configuration's loss, each with $L+1$ parameters. The additive model is the optimal first-order predictor. By Parseval's identity its mean-squared error equals the variance of $f$ left unexplained by per-layer effects, which we measure on full lattices, estimate out of sample at full-network scale, and report with every result as a certificate of how well any additive model can do. The coverage model itself is the second predictor. As allocators at matched memory, they attain the lowest KL divergence among the compared allocators on models from 30B to 355B parameters. Below four bits, the resulting allocations continue to solve code and reasoning tasks at budgets where allocations from gradient sensitivities no longer produce terminating generations.

Comments39 pages, 7 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑