arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向网络物理系统的、基于时间属性驱动的强化学习设计空间探索

Temporal Property-driven Design Space Exploration with Reinforcement Learning for Cyber-Physical Systems

Tagir Fabarisov, Maxime Cordy

arXiv 2608.23440首次发表:更新:

发表机构

SnT, University of Luxembourg(卢森堡大学 SnT)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对网络物理系统设计空间探索的难题,提出基于强化学习的时间属性驱动工作流,在矿用泵CPS案例中,其引导搜索需更少仿真次数即可找到最优设计。

AI 中文摘要

可配置网络物理系统(Cyber-Physical Systems, CPS)的设计空间探索需要可执行评估,因为设计选择会影响时序、故障传播、恢复行为及时间属性满足情况。重复随机执行使大型设计空间的穷举探索不切实际。本文提出一种基于时间属性驱动的CPS设计工作流,使用强化学习(Reinforcement Learning, RL)。设计阶段,RL智能体选择子系统备选方案以组装候选系统模型,该模型随后通过仿真评估,期间在线时间属性监视器观察运行时轨迹并生成功能属性违反指标。这些指标与评估的非功能项(预算、可恢复性、持续合规性及操作使用)结合,计算用于后续候选选择的奖励。该工作流在甲烷敏感型矿用泵CPS上评估,对应的可执行案例研究模型作为额外贡献提供。RL引导的搜索在26个回合(对应130次可执行仿真)后识别出实验中观测到的最高奖励设计,在相同可执行模型和奖励公式下,相比基于代理的贝叶斯优化和基于种群的遗传算法基线,达到这些设计所需的仿真次数更少。消融研究结果表明,基于价值的反馈和先前仿真轨迹的复用促成了这种减少。

英文摘要

Design-space exploration of configurable Cyber-Physical Systems (CPS) requires executable evaluation when design choices affect timing, fault propagation, recovery behavior, and temporal-property satisfaction. Repeated stochastic executions make exhaustive exploration impractical for large design spaces. This paper presents a temporal-property-driven CPS design workflow using Reinforcement Learning (RL). At design time, the RL agent selects subsystem alternatives to assemble a candidate system model. The model is then evaluated through simulation, during which online temporal-property monitors observe runtime traces and produce functional-property violation indicators. These indicators are combined with evaluated non-functional terms for budget, recoverability, sustained compliance, and operational use to calculate the reward used for subsequent candidate selection. The workflow is evaluated on a methane-sensitive mine-pump CPS. The corresponding executable case-study model is provided as additional contribution. RL-guided search identifies the highest-reward design observed in the experiments after 26 episodes (corresponds to 130 executable simulations). These designs were reached with fewer simulations than surrogate-guided Bayesian Optimization and population-based Genetic Algorithm baselines under the same executable model and reward formulation. Ablation study results indicate that value-based feedback and reuse of previous simulation traces contribute to this reduction.

CommentsAccepted manuscript for IECON 2026 - 52nd Annual Conference of the IEEE Industrial Electronics Society, Doha, Qatar, 18-21 October 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑