基于图的安全强化学习用于时变拓扑多智能体系统
Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying Topology
浏览论文内容
中文总结 AI 辅助
本文提出基于图的安全多智能体强化学习框架,通过控制障碍类函数筛选层确保安全,结合注意力Actor和GAT Critic,在真实机器人上验证了时变拓扑下的稳定性与安全性。
中文摘要 AI 辅助
本文提出了一种基于图的安全多智能体强化学习(MARL)框架,用于具有时变拓扑的协作导航。为了解决在具有感知约束的环境中确保安全性的关键挑战,通过控制障碍类函数(CBLF)动作筛选层引入了一种安全解耦机制。该机制弥合了离散LiDAR感知与连续安全约束之间的差距,确保无论学习进度如何,物理安全约束都得到严格满足。在此安全基础之上,提出了一种统一的结构架构,集成了基于注意力的Actor和基于图注意力网络(GAT)的集中式Critic。Actor利用值向量重构机制,通过协作跟踪误差矩阵显式编码相对几何关系,从而在时变通信拓扑下实现尺度不敏感的策略学习。同时,基于GAT的Critic对演化的交互结构进行建模,以实现准确的全局价值估计。所提出的框架在真实的差速驱动机器人平台上进行了验证,实验结果表明,在视野受限的动态场景中,该框架具有优越的稳定性和安全性。
英文摘要
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative navigation with time-varying topology. To address the critical challenge of ensuring safety in environments with sensing constraints, a safety-decoupled mechanism is introduced through a Control Barrier-Like Function (CBLF) action screening layer. This mechanism bridges the gap between discrete LiDAR perception and continuous safety constraints, ensuring that physical safety constraints are strictly satisfied regardless of the learning progress. Building upon this safety foundation, a unified structural architecture is proposed, integrating a attention-based actor and a Graph Attention Network (GAT) centralized critic. The actor utilizes a value vector reconstruction mechanism that explicitly encodes relative geometric relations through a collaborative tracking error matrix, enabling scale-insensitive policy learning under time-varying communication topologies. Meanwhile, the GAT-based critic models evolving interaction structures for accurate global value estimation. The proposed framework is validated on real differential-drive robot platforms, and experimental results demonstrate superior stability and safety in dynamic scenarios with limited fields-of-view.
发表机构
- School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University(北京交通大学机械与电子控制工程学院)
机构由 AI 辅助整理,请以论文原文为准。