arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19076cs.LGcs.AImath.STphysics.comp-phphysics.data-anstat.TH

双重下降是最小作用量原理

Double descent is the principle of least action

  • Penn Medicine Doylestown Hospital(宾夕法尼亚大学医学院多伊尔斯敦医院)

机构由 AI 辅助整理,请以论文原文为准。

Congzhou M Sha

AI总结:

该研究用统计力学解释双重下降现象:训练轨迹视为粒子在能量景观中游走,增加参数降低温度并增强权重正则化,从而在固定损失下降低测试误差。

AI中文摘要:

模型的测试误差随其参数数量 $d$ 的变化曲线先下降,在模型恰好能拟合训练数据时达到峰值,随后再次下降,展现出双重下降现象。我们用统计力学来解释这一现象。基于随机梯度的训练方法的训练轨迹是一个粒子在训练损失的能量景观中以诱导温度 $T$ 游走,而一个已平衡的运行会以玻尔兹曼分布给出的概率等可能地访问给定训练损失的每一个参数向量,这是统计力学的基本假设。由于训练从初始点开始且仅有有限时间进行扩散,它携带了有效的权重衰减,这使得每个参数都成为一个二次自由度。然后,均分定理将能量以 $T/2$ 的份额分配给这 $d$ 个自由度,因此在固定训练损失下,增加参数会降低温度,并将玻尔兹曼分布推向平稳路径。最后,增加参数只会降低平稳路径的 $L^2$ 范数,因此在固定损失下采样的解随着 $d$ 的增大而不太可能变大,这实际上增强了权重正则化。

英文摘要:

The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double descent phenomenon. We explain the phenomenon with statistical mechanics. The training trajectory of a stochastic gradient-based method is a particle wandering over the energy landscape of the training loss at an induced temperature $T$, and a run that has equilibrated visits every parameter vector of a given training loss equally often, the fundamental postulate of statistical mechanics, with probability given by the Boltzmann distribution. Because training starts at an initial point and has only finite time to diffuse, it carries an effective weight decay, which makes every parameter a quadratic degree of freedom. The equipartition theorem then distributes the energy among the $d$ degrees of freedom in shares of $T/2$, so at a fixed training loss adding parameters lowers the temperature and drives the Boltzmann distribution toward the stationary path. Finally, adding parameters can only lower the $L^2$ norm of the stationary path, so a solution sampled at fixed loss is less likely to be large with increasing $d$, effectively increasing weight regularization.

补充信息

↑