发表机构
Université Paris-Saclay; CNRS; ENS Paris-Saclay; Centre Borelli; Framatome(巴黎萨克雷大学; 法国国家科学研究中心; 巴黎萨克雷高等师范学院; 博雷利中心; 法马通公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对核电站负荷跟踪的行为克隆任务,提出将不同时间尺度变量编码至独立潜在空间的结构化架构,相比共享嵌入提升准确性与可行性,并用于热启动NMPC优化器,恢复完全可行性并降低约15%计算时间。
AI 中文摘要
用于工业控制的学习模型通常以整体准确率来评判,但组件层面的准确率并不能保证在嵌入其所服务的系统后具备安全性。我们在一个行为克隆任务上研究了这一差距:模仿专家非线性模型预测控制(NMPC)策略,以实现压水堆(PWR)的负荷跟踪,这是一个具有严格安全约束的工业系统。我们提出了一种结构化架构,将来自每个时间尺度的变量编码到独立的潜在空间中,反映系统的物理分解,然后在乘积潜在空间上训练控制器以模仿专家。在长时程滚动模拟中,与共享嵌入基线相比,分离的嵌入同时提高了准确性和可行性。敏感性分析进一步表明,我们的模型产生了与系统物理特性对齐的可解释表示。然而,无论架构如何,独立部署仍会留下几个百分比的轨迹不可行。使用我们的方法为NMPC优化器提供热启动,而非独立运行,我们恢复了完全的可行性和接近最优的成本,同时相对于专家控制器仍将计算时间削减了约15%,在突然的运行变化中甚至更多。
英文摘要
Learned models for industrial control are usually judged by aggregate accuracy, but accuracy at the component level does not guarantee safety once it is embedded in the system it is meant to serve. We study this gap on a behavior-cloning task: imitating an expert Nonlinear Model Predictive Control (NMPC) policy for load-following of a Pressurized Water Reactor (PWR), an industrial system with tight safety constraints. We propose a structured architecture encoding variables from each timescale into separate latent spaces, reflecting the physical decomposition of the system, before training a controller to imitate the expert on the product latent space. On long-horizon rollouts, separated embeddings improve both accuracy and feasibility compared with a shared-embedding baseline. Sensitivity analysis further shows that our model yields interpretable representations aligned with the system's physics. However, standalone deployment still leaves several percent of trajectories infeasible regardless of the architecture. Using our method to warmstart the NMPC optimizer rather than acting standalone, we recover full feasibility and near-optimal cost while still cutting computation time by $\sim$15% relative to the expert controller, and even more for abrupt operating changes.
Journal refNeurIPS 2026 Workshop AI Foundations for Power Grids, Dec 2026, Sydney (Australia), Australia