arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习何时细化:面向预算型神经算子PDE求解器的长时程强化学习

Learning When to Refine: Long-Horizon Reinforcement Learning for Budgeted Neural-Operator PDE Solvers

Ange Tong

arXiv 2610.06883首次发表:更新:

发表机构

National Research Tomsk State University(国立研究托木斯克理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对预算型神经算子PDE求解器,提出滚动验证策略改进(RV-PI),通过长时程强化学习决定细化时机,在浅水波和Brusselator基准上分别提升5.37%和2.31%的轨迹误差性能。

AI 中文摘要

神经算子为时间相关的偏微分方程提供了快速替代模型,但自回归部署产生了一个细化分配问题:预测误差随空间和时间变化,而沿轨迹只能投入有限数量的局部修正。我们将此问题表述为预算型自适应神经算子求解。一个全局傅里叶神经算子推进整个场,一个局部算子提出基于块的残差修正,一个集合感知选择器决定在何处细化。一个宏观策略决定何时以及花费多少剩余细化预算。我们引入了滚动验证策略改进(RV-PI),该方法通过实际延续滚动学习到的PDE求解器来评估可行的细化次数,将长时程优势转化为保守的策略目标,并且仅在保留轨迹误差改善时接受更新。在具有32次干预预算的浅水波基准上,RV-PI实现了三次种子平均轨迹相对L2误差为0.6910,比仅即时策略改进提高了5.37%,比RandomMacro提高了2.41%。在具有76次干预预算的强迫驱动Brusselator基准上,RV-PI达到了0.09954,比仅即时策略改进提高了2.31%,比RandomMacro提高了5.32%。这些结果表明,在固定细化预算下,局部修正的价值取决于其对自回归轨迹的下游影响,而不仅仅取决于其即时误差减少。

英文摘要

Neural operators provide fast surrogates for time-dependent PDEs, but autoregressive deployment creates a refinement-allocation problem: prediction errors vary over space and time, while only a finite number of local corrections can be committed along a trajectory. We formulate this as budgeted adaptive neural-operator solving. A global Fourier neural operator advances the full field, a local operator proposes patch-wise residual corrections, and a set-aware selector chooses where to refine. A macro policy decides when and how much of the remaining refinement budget to spend. We introduce rollout-verified policy improvement (RV-PI), which evaluates feasible refinement counts through actual continuation rollouts of the learned PDE solver, converts long-horizon advantages into conservative policy targets, and accepts an update only when held-out trajectory error improves. On the shallow-water benchmark with a 32-intervention budget, RV-PI achieves a three-seed mean trajectory relative L2 error of 0.6910, improving over immediate-only policy improvement by 5.37% and RandomMacro by 2.41%. On the forcing-driven Brusselator benchmark with a 76-intervention budget, RV-PI attains 0.09954, improving over immediate-only policy improvement by 2.31% and RandomMacro by 5.32%. These results show that, under a fixed refinement budget, the value of a local correction depends on its downstream effect on the autoregressive trajectory, not only on its immediate error reduction.

Comments21 pages, 4 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑