arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SMC-ES:形式验证控制策略的自动合成

SMC-ES: Automated synthesis of formally verified control policies

Riccardo Curcio, Toni Mancini, Enrico Tronci

arXiv 2607.15003首次发表:更新:

发表机构

Department of Computer Science, Sapienza University of Rome(罗马第一大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在安全关键环境中自动合成有形式保证的控制策略问题,提出基于仿真的新方法,开发SMC-ES算法,经实验验证其可行性,虽计算成本增加,但能提供形式保证且性能有竞争力。

AI 中文摘要

在安全关键环境中部署自主网络物理系统需要闭环控制策略,不仅性能良好,而且可证明安全可靠。基于学习的方法通常缺乏安全部署所需的形式保证。为此,我们提出一种基于仿真的新方法,能自动合成满足性能、安全和鲁棒性规范的形式保证策略。给定要验证的属性集、置信参数δ和允许的失败概率ε,该方法保证合成策略有证书,违反属性的概率至多为ε。我们开发了SMC-ES算法验证其可行性,并在一系列连续控制任务上进行评估,结果表明该算法虽计算成本增加,但能提供形式保证且性能有竞争力。

英文摘要

The deployment of autonomous cyber-physical systems in safety-critical environments requires closed-loop control strategies (i.e., policies) that are not only performant but also provably safe and robust. While learning-based methodologies such as Reinforcement Learning offer flexible and scalable approaches to automatically synthesize such controllers, they typically lack the formal guarantees necessary for safe deployment. To bridge this gap, we propose a novel simulation-based methodology to automatically synthesize policies with formal guarantees regarding performance, safety, and robustness specifications. Specifically, given a set of properties to verify, a confidence parameter $δ$ and an allowable failure probability $\varepsilon$, our method guarantees that the synthesized policy comes with a certificate: with confidence at least $1 - δ$, the probability of encountering a scenario where the given properties are violated is at most $\varepsilon$. We demonstrate the feasibility of our approach by developing SMC-ES, an algorithm that integrates Evolutionary Strategies with Statistical Model Checking-based verification. We evaluate SMC-ES on a suite of continuous control tasks using Gymnasium and Safety Gymnasium testbeds. Results show that, at the price of a sustainable increase in computational cost, our algorithm provides formal guarantees regarding performance, safety, and robustness specifications, while performing competitively against leading model-free Deep Reinforcement Learning (DRL) and Safe-DRL baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑