arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当神经代理无法加速求解器时:刚性耦合模拟中的运行时间占比、闭环漂移与不确定性门控的经济性

When a neural surrogate cannot accelerate a solver: runtime share, closed-loop drift, and the economics of uncertainty gating in a stiff coupled simulation

L. Thümmler, T. Kuroda

arXiv 2608.23075首次发表:更新:

发表机构

ETH Zürich; Max Planck Institute for Gravitational Physics (Albert Einstein Institute)(苏黎世联邦理工学院; 马克斯·普朗克引力物理研究所(阿尔伯特·爱因斯坦研究所))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现神经代理无法加速刚性耦合模拟中的隐式牛顿求解器,源于运行时间占比限制、离线精度失效及分布外门控失效等结构性障碍,相关结果为多物理场模拟的代理模型应用提供了关键警示。

AI 中文摘要

针对昂贵的内部求解器模块,学习得到的代理模型是实现多物理场模拟加速的常用途径。我们报告了一项受控的端到端负面结果,并识别出三个结构性障碍,这些障碍均与我们训练的网络性能无关。测试平台为隐式牛顿求解器,用于在广义相对论辐射流体动力学代码中耦合依赖能量的中微子辐射与物质,该求解器是每次调用中最昂贵的物理例程。首先,单次调用成本与运行时间占比是不同的量,仅后者对加速效果构成限制。独占自时间分析显示,目标模块占临界等级挂钟时间的16.9%,根据阿姆达尔定律,任何代理模型的加速上限约为1.2倍。一个单次调用成本低5.8倍的代理模型仅能与原求解器持平,而足够稳定以在无 fallback 机制下运行的配置也仅能达到同等性能。其次,离线精度无法对部署用代理模型进行排序:在14个网络中,汇集的斯皮尔曼误差与存活度的相关性(ρ=+0.73)是家族间混淆因素,在控制变量后该相关性消失(ρ=-0.04)。第三,正确的分布外门控无法加速离开其训练分布的循环。我们给出了盈亏平衡延迟比例的闭式表达式:由于访问的状态偏离数据流形73倍,门控会延迟96.8%至99.7%的单元,几乎与代理模型质量无关。包含自身成本后,门控循环的速度会降低0.94至0.96倍。我们进一步将稳定性与保真度分离:一次永不崩溃的门控运行在6000步内会累积线性的-19.9%密度偏差,该偏差是定向的、弹道式累积的偏差,而非自回归文献所关注的方差驱动发散。

英文摘要

Learned surrogates for expensive inner solver blocks are a widely pursued route to faster multiphysics simulation. We report a controlled, end-to-end negative result and identify three structural barriers, none of them a deficiency of the network we trained. The testbed is the implicit Newton solve coupling energy-dependent neutrino radiation to matter in a general-relativistic radiation-hydrodynamics code, its most expensive physics routine per call. First, per-call cost and share of runtime are different quantities, and only the second bounds acceleration. An exclusive self-time profile puts the target block at 16.9% of critical-rank wall clock, capping any surrogate at ~1.2x by Amdahl's law. A surrogate 5.8x cheaper per call merely ties the solver, and the configuration stable enough to run without fallback reaches only parity. Second, offline accuracy cannot rank surrogates for deployment: across fourteen networks the pooled Spearman error-versus-survival correlation (rho = +0.73) is a between-family confound that vanishes under control (rho = -0.04). Third, a correct out-of-distribution gate cannot accelerate a loop that leaves its training distribution. We give the break-even deferral fraction in closed form: because the visited states sit 73x off the data manifold, the gate defers 96.8 to 99.7% of cells, almost invariant to surrogate quality. Including its own cost, the gated loop is a 0.94 to 0.96x slowdown. We further separate stability from fidelity: a never-crashing gated run accumulates a linear -19.9% density bias over 6000 steps. The error is a directed, ballistically accumulating bias, not the variance-driven divergence the autoregressive literature targets.

Comments44 pages, 10 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑