arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

训练神经网络精度下限的深度定律:放大效应、残差缩放与量化感知训练悖论

Depth Laws for the Precision Floor of Trained Neural Networks: Amplification, Residual Scaling, and a Quantization-Aware Training Paradox

Ahmad S. Tarawneh

arXiv 2609.32060首次发表:更新:

AI 中文总结

本文提出深度定律,揭示神经网络精度下限由预测放大决定,随深度线性增长,并发现量化感知训练存在悖论:其增益随深度衰减,导致深度指数变陡。

AI 中文摘要

一个网络在精度崩溃之前需要多少比特,这个需求如何随深度增长?我们研究了在训练后量化(PTQ)和量化或噪声感知训练(QAT)下,MLP、CNN、Vision Transformers 以及九个预训练语言模型中的精度下限,即精度下降到随机水平一半时的扰动水平或位宽。(i)一阶理论通过一个全精度量,即预测放大 $G$,设定下限:$\eta_c=\Lambda/G$,且 $G^2$ 随深度线性增长,其速率与残差分支尺度的平方成正比。(ii)预测的深度指数等式 $\alpha_{PTQ}=\rho$ 在十三个训练架构中的十二个以及 GPT-2 从 12 层到 48 层中,在 95% 置信区间内成立,其中 $\Lambda=1.45\pm14\\%$ 跨越训练架构。(iii)残差分支按 $1/\sqrt{D}$ 缩放并采用预归一化可消除深度惩罚,每个量化器按步长规则将噪声转化为比特,给出 $b_c=(\alpha/\gamma)\log_2 D+C$。(iv)一个 QAT 悖论:噪声感知训练大致将浅层网络的可容忍噪声加倍,但增益随深度衰减,因此深度定律变陡($\alpha_{QAT}/\alpha_{PTQ}=1.45$-$1.47$,在两个数据集上各十个种子)。决策边界、跨层误差抵消和重尾并不决定下限。

英文摘要

How many bits does a network need before its accuracy collapses, and how does this grow with depth? We study the precision floor, the perturbation level or bit-width at which accuracy falls halfway to chance, in MLPs, CNNs, Vision Transformers and nine pretrained language models, under post-training quantization (PTQ) and quantization- or noise-aware training (QAT). (i) A first-order theory sets the floor through one full-precision quantity, the predictive amplification $G$: $η_c=Λ/G$, and $G^2$ grows linearly in depth at a rate proportional to the squared residual branch scale. (ii) The predicted equality $α_{PTQ}=ρ$ of depth exponents holds within 95% intervals in twelve of thirteen trained architectures and in GPT-2 from 12 to 48 layers, with $Λ=1.45\pm14\%$ across trained architectures. (iii) Residual branches scaled by $1/\sqrt{D}$ and pre-normalisation remove the depth penalty, and each quantizer turns noise into bits at a rate fixed by its step rule, giving $b_c=(α/γ)\log_2 D+C$. (iv) A QAT paradox: noise-aware training roughly doubles the tolerable noise of shallow networks, but the gain decays with depth, so the depth law steepens ($α_{QAT}/α_{PTQ}=1.45$-$1.47$ on two datasets, ten seeds each). Decision margins, cross-layer error cancellation and heavy tails do not set the floor.

Comments27 pages, 9 figures, 6 tables. Code: https://github.com/Ahmadtr/depth-laws-precision

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑