基于预测屏蔽的去中心化安全多智能体强化学习
Decentralized Safe Multi-Agent Reinforcement Learning via Predictive Shielding
浏览论文内容
中文总结 AI 辅助
提出一种去中心化框架,结合预测屏蔽与基于模型的有限时域Q学习,使多智能体在部署中安全调整策略,并引入无通信协议解决对称场景中的活锁问题。
中文摘要 AI 辅助
环境中越来越多地部署着多个机器人,它们执行独立任务,且对彼此的先验知识有限。部署此类多智能体系统带来了重大挑战。具体而言,与训练数据相比,部署状态的变化可能导致策略性能下降和安全性受损。虽然存在安全屏蔽来缓解这些风险,但它们通常是反应式的,这会在未见障碍物附近降低性能,并且是中心化的,限制了其可扩展性。为了解决这些问题,我们提出了一种去中心化框架,将预测屏蔽与基于模型的有限时域Q学习相结合。该方法使智能体能够在部署期间安全地调整其预训练策略。此外,为了缓解对称场景中的活锁问题,我们引入了一种无通信的冲突解决协议。
英文摘要
Environments are increasingly populated by multiple robots performing independent tasks with limited prior knowledge of each other. Deploying such multi-agent systems presents significant challenges. Specifically, shifts in deployment states compared to training data can lead to poor policy performance and compromised safety. While safety shields exist to mitigate these risks, they are typically reactive, which degrades performance near unseen obstacles,and centralized, limiting their scalability. To address this, we propose a decentralized framework that integrates predictive shielding with model-based finite horizon Q-learning. This approach allows agents to safely adapt their pre-trained policies during deployment. Furthermore, to mitigate livelocks in symmetric scenarios, we introduce a communication- free protocol for conflict resolution
发表机构
- ENSTA(国立高等先进技术学院)
- IP Paris(巴黎理工学院)
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。