arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29715cs.LGcs.AI

再验证优于有状态路由:面向分布漂移下的科学代理模型

Revalidation Beats Stateful Routing for Scientific Surrogates Under Distribution Shift

Harshil Lodhiya

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过流式基准RegimeShift-Surrogates比较再验证与有状态路由,发现基于当前窗口重新验证模型选择更优,平均对数遗憾更低,且优于所有有状态自适应方法。

中文摘要 AI 辅助

代理模型通常在开发阶段被选定,随后在新测量数据到来时保持不变。当噪声、输入支持范围或物理参数发生变化时,这种做法变得有风险。我们探究了此类变化是否需要一种有状态的自适应控制器,还是仅需在每个新批次上重新验证候选模型即可。为研究这一问题,我们构建了RegimeShift-Surrogates,一个可复现的流式基准,涵盖八个解析和动力学任务、四种平稳或漂移机制、十个保留种子以及八种经典、多层感知机和Kolmogorov-Arnold网络代理模型。确认性运行包含30,720次模型拟合和3,200个评分的部署窗口。在当前窗口中选择验证损失最低的模型,相对于逐窗口最优模型,平均对数遗憾为0.091;事后选择的最佳固定模型为0.192。配对差异为-0.101(层次自助法95%置信区间[-0.165, -0.040];Holm校正p=0.0469),在32个任务-场景组合中,再验证在26个中领先。所有有状态替代方案,包括指数平滑、双时间尺度自适应、Page-Hinkley重置或边际门控,均未改善汇总结果,而延迟偏差校正使其更差。最优选择也因任务而异:k近邻在阻尼振荡器上占优,普通KAN常被选用于二维曲面,MLP在Runge和Van der Pol任务上领先。在该基准中,新鲜的验证证据是有用的;而沿用旧证据往往无益。

英文摘要

Surrogate models are often chosen during development and then left in place as new measurements arrive. That practice becomes risky when noise, input support, or physical parameters change. We asked whether such changes call for a stateful adaptive controller, or whether it is enough to validate the candidate models again on each new batch. To study this question, we built RegimeShift-Surrogates, a reproducible streaming benchmark spanning eight analytic and dynamical tasks, four stationary or shifting regimes, ten held-out seeds, and eight classical, multilayer-perceptron, and Kolmogorov-Arnold network surrogates. The confirmatory run contains 30,720 model fits and 3,200 scored deployment windows. Choosing the model with the lowest validation loss in the current window yields mean log regret 0.091 against a per-window oracle; the best fixed model chosen in hindsight yields 0.192. The paired difference is -0.101 (hierarchical bootstrap 95% CI [-0.165, -0.040]; Holm-adjusted p = 0.0469), with revalidation ahead in 26 of 32 task-scenario combinations. None of the stateful alternatives, including exponential smoothing, dual-timescale adaptation, Page-Hinkley resets, or margin gating, improves the pooled result, and delayed bias correction makes it worse. Oracle choices also differ substantially by task: k-nearest neighbors dominate the damped oscillator, vanilla KAN is often selected for two-dimensional surfaces, and MLPs lead on the Runge and Van der Pol tasks. In this benchmark, fresh validation evidence is useful; carrying old evidence forward is often not.

发表机构

  • SlicedHealth

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑