AI 中文总结
该研究针对大参数下实解析位势的朗之万扩散,推导其在零集上的极限演化,发现演化偏向更高维奇异层,为随机梯度方法偏向泛化良好的奇异解提供了机制解释。
AI 中文摘要
我们考虑一般非负实解析位势 $V$ 和大参数 $\beta$ 对应的朗之万扩散过程 $dX_t = - \beta \nabla V(X_t) dt + \sqrt{2} dB_t$。当 $\beta$ 很大时,若过程从 $V$ 的零集出发,它会被限制在该零集内。我们推导了零集上的候选极限演化,为此根据称为局部学习系数的余维数局部度量及其重数将零集划分为层,进而证明与 $X$ 关联的狄利克雷形式在某种意义下收敛到对应于层上随机演化的狄利克雷形式层级,该演化强烈偏向更高维或“更奇异”的层。该结果由 Watanabe 奇异学习理论中关于过参数化统计模型的学习动力学及深度学习泛化谜题的问题所驱动,表明随机梯度方法倾向于偏向泛化性能良好的奇异解这一现象存在相关机制。
英文摘要
We consider the Langevin diffusion $dX_t = - β\nabla V(X_t) dt + \sqrt{2} dB_t$ for a general nonnegative real-analytic potential $V$ and a large parameter $β$. In the large-$β$ limit the process is confined to the zero set of $V$, assuming that it starts there. We derive a candidate limiting evolution on the zero set. To do so, the zero set is partitioned into strata according to a measure of local codimension known as the local learning coefficient and its multiplicity. It is then shown that the Dirichlet form associated with $X$ converges in a certain sense to a hierarchy of Dirichlet forms corresponding to a stochastic evolution on the strata. This evolution is strongly biased toward higher-dimensional, or "more singular", strata. This result is motivated by a question from Watanabe's singular learning theory regarding the learning dynamics of overparameterized statistical models and the generalization puzzle in deep learning. The result suggests a mechanism for the observation that stochastic gradient methods tend to be biased toward singular solutions that generalize well.