arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17758cs.ROcs.AI

CALOS:用于四旋翼安全深度强化学习的控制仿射李雅普诺夫流形安全层

CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors

发表机构博洛尼亚大学
查看机构详情
  • University of Bologna(博洛尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Fabrizio Cesareo, Sebastiano Mengozzi, Nicola Mimmo, Andrea Acquaviva

首次发表
浏览论文内容

中文总结 AI 辅助

CALOS提出一种基于控制仿射李雅普诺夫函数的运行时安全层,通过二次规划修正四旋翼策略输出,在Isaac Lab中实现零姿态违规并将横向误差降低55-60%,同时加速训练收敛。

中文摘要 AI 辅助

深度强化学习在四旋翼控制中展现了卓越的能力,然而学习到的策略在训练或部署期间无法保证遵守安全约束。我们提出了CALOS(控制仿射李雅普诺夫流形安全),一种运行时安全层,在不修改底层学习算法的情况下对四旋翼施加姿态约束。CALOS将四个倾斜角不等式和李雅普诺夫下降条件表述为单个二次规划,其解是对策略标称扭矩输出的最小范数修正。该二次规划通过三维扭矩空间上的活动集枚举精确求解,其计算成本足够低,可在数千个并行仿真环境中实时执行约束,满足现代大规模并行深度强化学习训练的要求。在NVIDIA Isaac Lab中的轨迹跟踪任务评估中,与无约束的近端策略优化基线相比,CALOS将横向跟踪误差降低了55-60%,同时在训练轨迹上实现了零姿态约束违反。通过将探索限制在状态空间的安全区域,该安全层还加速了训练收敛并提高了数据效率,而不会产生次优策略。

英文摘要

Deep Reinforcement Learning has demonstrated remarkable capability in quadrotor control, yet learned policies offer no guarantee of respecting safety constraints during training or deployment. We present CALOS (Control-Affine Lyapunov On-manifold Safety), a runtime safety layer that enforces attitude constraints on a quadrotor without modifying the underlying learning algorithm. CALOS formulates four tilt-angle inequalities and a Lyapunov descent condition as a single quadratic program whose solution is the minimum-norm correction to the nominal torque output of the policy. The quadratic program is solved exactly via active-set enumeration over the three-dimensional torque space, with a computational cost low enough to enforce constraints in real time across thousands of parallel simulation environments, as required by modern massively parallel Deep Reinforcement Learning training. Evaluated on trajectory-tracking tasks in NVIDIA Isaac Lab, CALOS reduces lateral tracking error by 55-60% relative to an unconstrained Proximal Policy Optimization baseline while achieving zero attitude-constraint violations on the training trajectory. By restricting exploration to safe regions of the state space, the safety layer also accelerates training convergence and improves data efficiency without producing suboptimal policies.

↑