LyEvO:用于安全且鲁棒的现实迁移策略学习的李雅普诺夫引导进化优化算法
Safe and Robust Neural Policy Learning with Statistical Verification for Sim-to-Real Deployment in Robotics
- Sapienza University of Rome(罗马大学)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出LyEvO框架,结合约束进化优化、统计模型检查与李雅普诺夫分析,用于安全鲁棒的现实迁移策略学习,经Cartpole和3D四旋翼实验验证有效。
AI中文摘要:
在仿真环境中训练安全且鲁棒的控制器,并系统评估其对现实部署的就绪程度,仍是现实迁移(sim-to-real)领域的关键挑战。为解决该问题,本文提出LyEvO——一种基于物理原理的框架,它结合了约束进化优化(constrained Evolutionary Optimization)、基于统计模型检查(Statistical Model Checking, SMC)的验证方法,以及李雅普诺夫(Lyapunov)稳定性分析。LyEvO利用系统动力学的先验知识,通过李雅普诺夫分析计算初始候选稳定区域;随后的迭代循环会使用该区域内的运行场景,联合优化策略并进行统计验证,再根据验证结果扩展区域边界。该集成流程为评估部署就绪程度提供了实用标准。本文在Cartpole和3D四旋翼(3D Quadrotor)基准上,通过大量仿真和针对性现实实验对LyEvO进行评估,验证了其安全且鲁棒的现实迁移能力。
英文摘要:
Synthesizing safe and robust neural controllers in simulation for reliable sim-to-real deployment remains a critical challenge in robotics. Existing learning-based methods typically lack safety and performance guarantees over an explicitly defined operating region, while post-training verification techniques provide no mechanism to refine controllers when safety violations are detected. To bridge this gap, we propose a curriculum-driven framework that tightly integrates scenario-based Evolution Strategy with Statistical Model Checking-based verification in a closed-loop procedure. Starting from a candidate region, our approach co-optimizes policy performance while progressively enlarging its safe operating boundaries. Upon termination, it yields a neural controller together with a region over which safety and performance are statistically verified. Extensive evaluations on Cartpole and 3D Quadrotor benchmarks, showing 6.14x and 224.04x expansions, respectively, of the safe operating region over mathematically certified ones, together with physical experiments under both nominal conditions and severe dynamic perturbations, demonstrate that our learned controllers consistently outperform established control-theoretic and learning-based baselines. Furthermore, we show that the size of the verified region serves as a quantitative indicator of policy quality before deployment. These results establish our framework as an automated pipeline for learning, assessing and deploying safe and robust neural controllers from simulation to reality.