AI 中文总结
该研究证明ReLU训练中,梯度下降的离散导数与梯度流极限的导数存在偏差,揭示了奇异曲率相关的微分一致性问题,为相关训练方法提供了理论分析基础。
AI 中文摘要
梯度下降(GD)是梯度流的显式欧拉格式,但状态精确的连续时间替代模型在微分后未必仍保持精确。对于每个固定的非共振步长,普通自动微分可精确对已执行的硬ReLU GD程序求导。我们证明,在固定有限时间范围内,GD状态收敛,且这些精确离散导数趋近于无事件的区域传播子,而极限流的导数还包含速度归一化的激活事件转移。前点斯蒂尔杰斯表示将绝对连续的区域海森矩阵与原子界面曲率分离;一个非零梯度跳跃会产生恰好秩1的端点偏差,且当事件严格时,全局凸性会阻止多事件完全抵消。不过,一类标准的全局1-强凸残差ReLU平方损失模型,在开初始化集上会呈现任意大的倒数灵敏度比,且具有一致横截距裕度。相同的离散-流分解可扩展至参数和反向模式伴随;在标量或自主正态区域中,通过解析平滑及一致事件定位可恢复流灵敏度。这些结果针对确定性全批量、有限时间范围的动力学,具有稳定的有限行程的同向分离横向事件;它们是一致性定理,而非大规模训练的普遍性断言。
英文摘要
Outer-learning algorithms use infinitesimal sensitivities to propose finite changes to initialization or training parameters. For hard-ReLU training, the derivative of the finite program and the derivative of its flow limit do not by themselves specify the response at a chosen radius. We characterize the intervening regime in which the perturbation radius is proportional to the GD step. Integer event rounding then survives at leading order: smooth Euler bias shifts each discrete phase, and upstream rounding moves downstream branch boundaries. We derive the crossing indices and a uniform endpoint expansion for finitely many separated transverse events in piecewise-$C^2$ dynamics, away from recursive phase boundaries. In contractive affine regions, an explicit remainder and complete branch verification certify finite candidate comparisons. Scalar phase frequencies and a coupled feedback ablation test the mechanism; frozen nonlinear-network experiments show radius-dependent prediction accuracy, including incomplete branch matches and failed-word tails. Together with local AD and uniform flow consistency, the result identifies sufficient response regimes: differentiating training is a choice of perturbation resolution as well as a choice of derivative.