arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

持续图多智能体强化学习

Continual Graph Multi-Agent Reinforcement Learning

Tommaso Marzi, Ahmed Hendawy, Jan Peters, Carlo D'Eramo, Andrea Cini, Cesare Alippi

arXiv 2610.10302首次发表:更新:

发表机构

IDSIA USI-SUPSI, Università della Svizzera italiana; Technical University of Darmstadt; Robotics Institute Germany (RIG); German Research Center for AI (DFKI); University of Würzburg; EPFL; Politecnico di Milano(IDSIA USI-SUPSI,瑞士意大利语大学; 达姆施塔特工业大学; 德国机器人研究所(RIG); 德国人工智能研究中心(DFKI); 维尔茨堡大学; 洛桑联邦理工学院; 米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CGMARL框架,将任务映射为属性图以利用结构信息,并推出GRAFO基准及FROG方法,通过冻结图主干缓解遗忘,显著提升持续多智能体强化学习性能。

AI 中文摘要

在持续多智能体强化学习(CMARL)中,智能体跨任务序列学习协作策略,旨在有效适应新任务的同时保持解决先前遇到任务的能力。在许多应用中,任务在底层结构上有所不同,这些结构可以代表例如不同的运行条件或目标配置(如电网中的不同网络拓扑或编队控制中的不同排列)。现有的CMARL方法缺乏专用机制来在学习新任务时利用这种结构信息,未能促进迁移并缓解遗忘。为填补这一空白,我们提出持续图多智能体强化学习(CGMARL),这是一个针对CMARL问题的新框架,其中任务序列被映射为一系列属性图,每个图建模特定于任务的结构。在CGMARL中,每个图决定相应任务的环境动态(下一状态和/或奖励)以及智能体数量。然后,我们提出基于图的编队(GRAFO),这是第一个CGMARL基准,并展示了遗忘如何在此设置中出现。最后,为解决这一局限,我们提出冻结图编码器(FROG),一种依赖冻结图主干来保留基于图的CMARL策略中过去结构信息的方法。在GRAFO上的实验表明,将FROG与现有持续学习方法配对可显著提升多个CGMARL场景的性能。

英文摘要

In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In many applications, tasks differ in their underlying structure, which can represent, for example, distinct operational conditions or target configurations (e.g., different network topologies in power grids or arrangements in formation control). Existing CMARL methods lack dedicated mechanisms to leverage this structural information when learning new tasks, failing to promote transfer and mitigate forgetting. To fill this gap, we propose Continual Graph Multi-Agent Reinforcement Learning (CGMARL), a novel framework for CMARL problems in which task sequences are mapped into a series of attributed graphs, each modeling a task-specific structure. In CGMARL, each graph determines the environment dynamics (next states and/or rewards) and the number of agents for the corresponding task. Then, we present Graph-based Formation (GRAFO), the first CGMARL benchmark, and show how forgetting arises in this setting. Finally, to address this limitation, we propose Frozen Graph Encoder (FROG), a method that relies on a frozen graph backbone to preserve past structural information in graph-based CMARL policies. Experiments on GRAFO show that pairing FROG with existing CL methods substantially improves performance on multiple CGMARL scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑