arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应龙格 - 库塔步长控制带来的是训练损失,而非泛化能力:关于 RK - Adam 优化器的公平计算匹配研究

Adaptive Runge-Kutta Step Control Buys Training Loss, Not Generalization: An Honest Compute-Matched Study of RK-Adam Optimizers

Akhilesh Gogikar

arXiv 2607.14516首次发表:更新:

AI 中文总结

研究将高阶 RK 积分器用于神经网络优化器,构建 Adam 变体并在计算匹配协议下评估,发现其“适应性”虚幻,修复后训练损失降低但对初始步长敏感,未达测试精度,还证实梯度平均是隐式正则化器,高阶自适应积分优势不明显。

AI 中文摘要

将优化器解释为梯度流离散化促使人们将高阶龙格 - 库塔(RK)积分器应用于神经网络。我们构建了一个具有代表性的 Adam 变体(博加茨基 - 尚皮内 3(2) RK 对、FSAL 重用、局部误差步长控制),并在严格的计算匹配协议下对其进行评估,即给每种方法相同的梯度评估预算,而这在该领域文献中很少被执行。在此协议下,RK 变体在小批量和全批量(RK 的最佳情况)训练中的训练损失都比普通 Adam 差。对其进行分析表明,“适应性”是虚幻的:归一化误差远低于容差,步长从第一步就固定在其增长上限(98 - 100%的步数),并且没有 rtol x hmax x h0 设置能使其起作用;跨越 100 倍的容差给出的轨迹完全相同。该方法实际上是以 3 - 4 倍成本使用平均梯度的固定步长 Adam。修复它(真正的拒绝分支;应用映射上的误差)会反转全批量结果——训练损失比调优后的 Adam 低约 40 倍——并且固定步长控制将适应性(一种出现的热身和增长调度)分离为机制。但这种增益对初始步长很脆弱,且未达到测试精度。一项预先注册的后续研究排除了明显的解释:更深层次的最小化不会过拟合,明确的温度旋钮只会有害——留下一种轨迹效应,控制器选择的最小值在相同深度下比一阶下降泛化能力低 1.3 - 3.4 个点。一项 n = 10 的研究证实了一个次要效应:梯度平均是一种真正的隐式正则化器,在 10/10 个种子上击败了学习率匹配的 Adam 和 AdamW——然而 RMSprop 和 NAdam 以三分之一的每步成本与之匹配或超越它。高阶自适应积分带来了更深层次的确定性最小化和小的正则化效果,但更便宜、调优良好的一阶基线已经能提供这些。

英文摘要

Interpreting optimizers as gradient-flow discretizations has motivated applying higher-order Runge-Kutta (RK) integrators to neural networks. We build a representative Adam variant (Bogacki-Shampine 3(2) RK pair, FSAL reuse, local-error step control) and evaluate it under a strict compute-matched protocol giving every method the same gradient-evaluation budget - an accounting this literature rarely enforces. Under it the RK variant loses to plain Adam on training loss in both minibatch and full-batch (RK's best-case) training. Instrumenting it shows the "adaptivity" is illusory: normalized error stays far below tolerance, the step size pins at its growth cap from step one (98-100 percent of steps), and no rtol x hmax x h0 setting makes it act; tolerances spanning 100x give bit-identical trajectories. The method is exactly fixed-step Adam with an averaged gradient at 3-4x cost. Repairing it (true reject branch; error on the applied map) reverses the full-batch result - about 40x lower training loss than tuned Adam - and a fixed-step control isolates adaptivity (an emergent warmup-and-growth schedule) as the mechanism. But the gain is fragile to the initial step size and does not reach test accuracy. A pre-registered follow-up rules out the obvious explanations: deeper minimization does not overfit, and an explicit temperature knob only hurts - leaving a trajectory effect, the controller selecting a minimum generalizing 1.3-3.4 points below first-order descent at equal depth. An n=10 study confirms one secondary effect: gradient averaging is a genuine implicit regularizer, beating lr-matched Adam and AdamW on 10/10 seeds - yet RMSprop and NAdam match or beat it at a third the per-step cost. Higher-order adaptive integration buys deeper deterministic minimization and a small regularization effect, but nothing a cheaper, well-tuned first-order baseline does not already provide.

Comments10 pages, 4 figures. Code, logs, and result JSONs: https://github.com/akigogikar/RK4Optimizer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑