arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

与未来协作者合作:交错参与下的多智能体强化学习

Cooperating with Future Collaborators: Multi-Agent RL under Staggered Participation

Jianglin Qiao, Siyi Hu, Thien Hoang Nguyen, Zehong Cao, Salah Sukkarieh

arXiv 2610.07578首次发表:更新:

发表机构

ACFR, The University of Sydney; Curtin University; Adelaide University(悉尼大学ACFR; 科廷大学; 阿德莱德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体强化学习中的交错参与问题,提出SPL训练增强方法,通过前瞻性获取监督和结果监督接收者学习,在60个设置中平均提升任务完成度14.1%,并扩展至八智能体及具身环境。

AI 中文摘要

在合作式多智能体强化学习(MARL)中,智能体通常在并发参与下进行训练,而在许多任务中,一些智能体更早行动并留下与任务相关的信息,这些信息对后来参与的智能体有用。我们将这种设置研究为交错参与(SP),它引入了跨时间、跨智能体的学习依赖,因为早期行动可能通过其提供的信息以及使用该信息的后续策略来影响回报。因此,在SP下学习既需要识别哪些信息对未来决策有用,也需要学习后来的智能体应如何使用这些信息。我们提出了交错参与学习(SPL),这是一种训练时增强方法,通过为早期智能体提供前瞻性获取监督和为后期智能体提供结果监督的接收者学习来解决这两个部分。我们在多种基于策略的MARL骨干网络、环境和交错参与模式上评估了SPL。在60个MPE/RWARE骨干设置比较中,SPL在每个案例中都取得了更高的观测平均任务完成度,平均差异为14.1%。这些收益还扩展到八智能体团队以及Isaac Lab中基于物理的UAV-UGV环境,提供了跨算法、时间和具身设置的证据。

英文摘要

In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents participating later. We study this setting as staggered participation (SP), which introduces a cross-time, cross-agent learning dependency because an early action may affect the return through the information it provides and the later policy that uses it. Learning under SP therefore requires both identifying what information is useful for future decisions and learning how later agents should use it. We propose Staggered Participation Learning (SPL), a training-time augmentation that addresses these two parts with prospective acquisition supervision for earlier agents and outcome-supervised receiver learning for later agents. We evaluate SPL across multiple policy-based MARL backbones, environments, and staggered-participation patterns. Across 60 MPE/RWARE backbone setting comparisons, SPL achieves higher observed mean task completion in every case, with an average difference of 14.1%. The gains also extend to eight-agent teams and a physics-based UAV-UGV environment in Isaac Lab, providing evidence across algorithmic, temporal, and embodied settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑