arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2412.00534cs.LGcs.AIcs.MA

迈向多智能体强化学习中的容错性

Towards Fault Tolerance in Multi-Agent Reinforcement Learning

  • Tsinghua University(清华大学)
  • QiYuan Lab(启元实验室)
  • Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University(北京信息科学与技术国家研究中心(BNRist),清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuchen Shi, Huaxin Pei, Liang Feng, Yi Zhang, Danya Yao

更新

AI总结:

针对多智能体强化学习中故障导致的状态混乱与样本不平衡问题,本文提出注意力 actor-critic 架构和优先级采样策略,并开源高度解耦的容错 MARL 平台。

AI中文摘要:

智能体故障对多智能体强化学习(MARL)算法的性能构成重大威胁,并带来两个关键挑战。首先,智能体往往难以从意外故障造成的混乱状态空间中提取关键信息。其次,回放缓冲区中在故障前后记录的转移会不均衡地影响训练,从而导致样本不平衡问题。为克服这些挑战,本文通过将优化的模型架构与定制的训练数据采样策略相结合,增强 MARL 的容错性。具体而言,在 actor 和 critic 网络中引入注意力机制,以自动检测故障并动态调节给予故障智能体的注意力。此外,本文引入一种优先级机制,以选择性地采样对当前训练需求至关重要的转移。为进一步支持该领域的研究,我们设计并开源了一个高度解耦的容错 MARL 代码平台,旨在提高研究相关问题的效率。实验结果证明了我们的方法在处理各种类型的故障、发生在任意智能体上的故障以及在随机时间出现的故障方面的有效性。

英文摘要:

Agent faults pose a significant threat to the performance of multi-agent reinforcement learning (MARL) algorithms, introducing two key challenges. First, agents often struggle to extract critical information from the chaotic state space created by unexpected faults. Second, transitions recorded before and after faults in the replay buffer affect training unevenly, leading to a sample imbalance problem. To overcome these challenges, this paper enhances the fault tolerance of MARL by combining optimized model architecture with a tailored training data sampling strategy. Specifically, an attention mechanism is incorporated into the actor and critic networks to automatically detect faults and dynamically regulate the attention given to faulty agents. Additionally, a prioritization mechanism is introduced to selectively sample transitions critical to current training needs. To further support research in this area, we design and open-source a highly decoupled code platform for fault-tolerant MARL, aimed at improving the efficiency of studying related problems. Experimental results demonstrate the effectiveness of our method in handling various types of faults, faults occurring in any agent, and faults arising at random times.

补充信息

↑