arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向低保守性的安全强化学习:无人机飞行控制案例研究

Towards Safe Reinforcement Learning with Reduced Conservativeness: A Case Study on Drone Flight Control

Loizos Hadjiloizou, Michael C. Welle, Hang Yin, Danica Kragic

arXiv 2608.26852首次发表:更新:

AI 中文总结

该研究提出一种降低形式化方法保守性的安全强化学习框架,通过zonotopic可达性分析保障实时安全,经无人机峡谷飞行验证,可在确保安全的同时提升控制器在线训练的探索性。

AI 中文摘要

将形式化方法融入强化学习(RL)有望实现两全其美,兼具形式化保障的鲁棒性与RL的适应性及学习能力,但需精心设计以平衡安全性与探索性。本研究提出一种框架,可在确保系统安全的同时缓解探索性损失。具体而言,引入一种限制性更低的方法,通过利用在线采集的数据优化扰动模型,降低形式化方法的保守性;采用计算高效的zonotopic可达性分析对基于学习的控制器进行安全性评估,以支持实时实现。我们在无人机穿越峡谷的真实飞行场景中验证该框架,无人机受未知外部扰动影响,框架需在线学习这些扰动并相应调整安全保障。结果表明,该框架可在不损害系统安全的前提下,实现基于学习的控制器的限制性更低的在线训练。

英文摘要

Incorporating formal methods into reinforcement learning (RL) has the potential to result in the best of both worlds, combining the robustness of formal guarantees with the adaptability and learning capabilities of RL, though careful design is needed to balance safety and exploration. In this work, we propose a framework to mitigate this loss of exploration while still allowing for the safety of the system to be ensured. Specifically, we introduce a less restrictive method that can reduce the conservativeness of formal methods by refining a disturbance model using online collected data and it evaluates the safety of a learning-based controller, using computationally efficient zonotopic reachability analysis for the safety analysis to facilitate a real-time implementation. We validate the framework in a real-world drone flight through a canyon, where the drone is subjected to unknown external disturbances and the framework is tasked with learning those disturbances online and adjusting the safety guarantees accordingly. The results show that the framework enables a less restrictive online training of learning-based controllers without compromising the safety of the system.

Comments7 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑