arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13204cs.LGcs.CL

GradRepair-ODE:神经ODE训练的可认证梯度修复

GradRepair-ODE: Certified Gradient Repair for Neural ODE Training

  • Purdue University(普渡大学)
  • Rice University(莱斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Ziqian Bi, Xin Liang Chia

AI总结:

针对神经ODE训练中数值求解器导致梯度不可靠的问题,提出GradRepair-ODE框架,通过检查、修复和拒绝机制认证梯度,在合成系统中将不安全步骤降至零。

AI中文摘要:

神经常微分方程在训练循环内部使用数值求解器。求解器决定前向轨迹,同时也影响传递给优化器的梯度。这种耦合给科学机器学习和连续时间生成建模(包括扩散概率流常微分方程和流匹配模型)带来了可靠性问题。在宽松的步长、刚性动力学、混沌敏感性或事件不连续条件下,可微分的ODE流水线可能返回方向在数值上可疑的有限梯度。我们提出GradRepair-ODE,一个在优化器步骤中检查、修复和拒绝ODE梯度的可靠性框架。该方法计算多个梯度候选,通过方向有限差分检查和求解器诊断进行比较,诊断可能的数值失效模式,通过路径切换或更严格的重计算修复选定的梯度,并拒绝下降方向无法认证的步骤。在六个合成ODE系统中,GradRepair-ODE保持低风险系统不变,将Robertson和Lorenz梯度修复到与严格参考的余弦相似度为1.000,将不安全的接受步骤从37减少到0,并拒绝事件不连续案例而不是应用未认证的更新。本文主张训练契约的一个简单改变:ODE梯度应附带数值证据到达优化器。

英文摘要:

Neural ordinary differential equations use numerical solvers inside the training loop. The solver determines the forward trajectory and also affects the gradient passed to the optimizer. That coupling creates a reliability problem for scientific machine learning and continuous-time generative modeling, including diffusion probability-flow ordinary differential equations and flow-matching models. Under loose step sizes, stiff dynamics, chaotic sensitivity, or event discontinuities, a differentiable ODE pipeline can return a finite gradient whose direction is numerically suspect. We introduce GradRepair-ODE, a reliability framework for checking, repairing, and rejecting ODE gradients at the optimizer step. The method computes several gradient candidates, compares them with directional finite-difference checks and solver diagnostics, diagnoses likely numerical failure modes, repairs selected gradients through path switching or stricter recomputation, and rejects steps whose descent direction cannot be certified. In six synthetic ODE systems, GradRepair-ODE leaves low-risk systems unchanged, repairs Robertson and Lorenz gradients to cosine similarity 1.000 against a strict reference, reduces unsafe accepted steps from 37 to 0, and rejects an event-discontinuous case instead of applying an uncertified update. The paper argues for a simple change in the training contract: an ODE gradient should reach the optimizer with numerical evidence attached.

补充信息

↑