arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SRL-MPC:形状感知强化学习模型预测控制

SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

Ruihua Han, Rui Gao, Zhe Liu, Xinyi Wang, Chang Chen, Shuai Wang, Qi Hao, Jia Pan, Hengshuang Zhao

arXiv 2608.21175首次发表:更新:

发表机构

The University of Hong Kong; Southern University of Science and Technology; University of Michigan; Shenzhen Institutes of Advanced Technology(香港大学; 南方科技大学; 密歇根大学; 深圳先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对异构人群与机器人集群的形状感知导航难题,提出SRL-MPC方法,结合MPC的安全泛化性与RL的适应性,实验显示其在安全性、适应性上优于基准方法。

AI 中文摘要

在异构人群与机器人集群中实现安全高效的形状感知导航仍是一项挑战。传统方法常假设机器人同质、工作空间稀疏、几何形状简化、采用离线计算或手工设计参数以简化问题,这限制了其在密集人群场景中的部署。为此,我们提出形状感知强化学习模型预测控制(Shape-Aware Reinforcement Learned Model Predictive Control,SRL-MPC),一种无需几何简化即可在异构形状人群中实现安全、高效、自适应导航的方法。为编码形状感知安全性,我们基于支撑函数变换从几何分离特征(Geometric Separation Features,GSFs)构建高阶控制障碍函数(High-Order Control Barrier Function,HOCBF)约束。随后,强化学习(Reinforcement Learning,RL)框架学习一种神经策略,该策略读取GSF并输出实时MPC参数更新,使MPC求解器能够适配邻近人群的几何形状。SRL-MPC的核心优势在于,它在保留MPC的安全结构与泛化性的同时,整合了RL的适应性与智能性。在含任意形状机器人集群的随机人群场景中开展的实验验证了SRL-MPC的有效性、可扩展性与鲁棒性。结果表明,SRL-MPC在安全性与适应性上显著优于代表性基准方法。项目网站:this https URL

英文摘要

Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape-aware safety, we formulate high-order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real-time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL-MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL-MPC. The results show that SRL-MPC substantially outperforms representative baselines in safety and adaptability. Project website: https://hanruihua.github.io/srl_mpc_project/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑