arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11475cs.LGmath.OCstat.ML

罕见门分歧会限制可塑性:梯度流何时会错误预测有限批次SGD

Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD

Ruoyu Zhao, Mingxuan Zhang, Jianbo Dai, Jiaqi Wu, Chenyu Zhu, Tong Che

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现总体梯度流会错误预测有限批次SGD,在双单元ReLU回归任务中,小步长与长预训练的联合极限会导致在线SGD失效,需满足特定样本与批次大小条件才能恢复。

中文摘要 AI 辅助

总体梯度流是用于推理神经网络如何适应(包括预训练后)的常用工具。本文表明,它会在定性层面错误预测有限批次随机梯度下降(SGD),并将这种差异追溯到特定机制。在双单元ReLU回归中,源任务会驱动两个神经元趋于正比例,而目标任务则奖励将它们分开。经过时长为T的源训练后,梯度流会在与T成线性关系的时间内恢复目标任务的性能。然而,在两个阶段中使用批次大小b和步长η的在线SGD,一旦T≳log(b/η),在量级为e^(c/η)的整个时间范围内,就会以高概率失效,且在具有高斯概率超过1%的显式初始化集合上是一致的。对于每个固定的T,小步长SGD仍能恢复,因此失效需要小步长和长预训练的联合极限。在目标克隆处,总体不稳定性完全由两个ReLU门存在分歧的输入所承载。对于角度为δ的单元,这些输入形成概率为δ/π的楔形,且权重衰减会在预训练期间将角度指数缩小。在其他所有输入上,两个单元都会收到相同的随机线性更新,这会在条件期望中缩小它们的分离度。通过沿精确在线递归对采样楔形的累积概率进行边界限定(不使用扩散近似),可得出:从相同的源-梯度流检查点以固定概率恢复,在e^(c/η)次更新内,需要Nb≳e^(λT)个目标样本,且批次大小b≳ηe^(λT),其中N为更新次数,λ为权重衰减系数。在模拟中,恢复情况近似为分歧预算bδ/η的函数,并在该时间范围内达到饱和。

英文摘要

Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining. We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism. In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them. After source training for time $T$, gradient flow recovers on the target in time linear in $T$. Online SGD with batch size $b$ and step size $η$ in both phases instead fails with high probability throughout a horizon of order $e^{c/η}$ once $T \gtrsim \log(b/η)$, uniformly on an explicit set of initializations with Gaussian probability above one percent. For each fixed $T$, small-step SGD still recovers, so the failure requires the joint limit of small steps and long pretraining. At the target clone, the population instability is carried entirely by inputs on which the two ReLU gates disagree. For units at angle $δ$ these inputs form a wedge of probability $δ/π$, and weight decay shrinks the angle exponentially during pretraining. On every other input both units receive the same random linear update, which contracts their separation in conditional expectation. Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within $e^{c/η}$ updates, requires $Nb \gtrsim e^{λT}$ target samples and batch size $b \gtrsim ηe^{λT}$, where $N$ counts updates and $λ$ is the weight decay. In simulations, recovery is approximately a function of the disagreement budget $bδ/η$ and saturates in the horizon.

发表机构

  • City University of Hong Kong(香港城市大学)
  • Microsoft(微软公司)
  • NVIDIA Research(英伟达研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑