arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03667cs.LGcs.AI

离线多智能体强化学习中基于序列模型的分布外泛化

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

发表机构InstaDeep公司 · 非洲数学科学研究所 · 斯坦陵布什大学
查看机构详情
  • InstaDeep(InstaDeep公司)
  • AIMS(非洲数学科学研究所)
  • Stellenbosch University(斯坦陵布什大学)

机构由 AI 辅助整理,请以论文原文为准。

Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob, Siddarth Singh, Juan Claude Formanek, Felix Chalumeau, Omayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, … 展开作者

Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob, Siddarth Singh, Juan Claude Formanek, Felix Chalumeau, Omayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi, Louay Ben Nessir, Simon Verster Du Toit, Daniel Rajaonarivonivelomanantsoa, Asim Awad Osman, Arnol Manuel Fokam, Refiloe Shabe, Arnu Pretorius

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对离线MARL的分布外泛化挑战,扩展序列建模架构适配多任务设置,发现任务多样性缩放是零样本迁移关键,多任务方法在四环境测试中较单任务模型提升3.2倍且优于基线。

中文摘要 AI 辅助

对未见任务的泛化仍是离线多智能体强化学习(MARL)中的核心挑战。本研究对离线场景下的零样本任务泛化进行了原则性分析,并针对任务多样性、数据集规模与网络容量相关的缩放行为开展了广泛实证研究。为支撑该研究,我们扩展了离线序列建模架构,使其能够处理多任务观测空间与动作空间,以及不同任务间可变的智能体数量。核心发现是:缩放任务多样性而非单纯的数据集规模,是实现稳健零样本迁移的主导因素。通过在Connector、RWARE、SMAX与LBF四个具有挑战性的环境中开展大规模实验,我们证明,与单任务模型相比,我们的多任务方法在保留的测试任务上实现了3.2倍的平均提升,且始终优于强大的行为克隆基线。这些结果表明,可泛化MARL智能体的开发应优先考虑训练分布的多样性,包括不同数量的智能体,为有效扩展离线MARL提供了路线图。

英文摘要

Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement learning (MARL). In this work, we present a principled analysis of zero-shot task generalisation in the offline setting and conduct an extensive empirical investigation into the scaling behaviour governing task diversity, dataset size, and network capacity. To facilitate this study, we extend offline sequence modelling architectures to handle multi-task observation and action spaces alongside variable agent counts across tasks. Our primary finding is that scaling task diversity---rather than sheer dataset size is the dominant factor in achieving robust zero-shot transfer. Through large-scale experiments across four challenging environments (Connector, RWARE, SMAX, and LBF), we demonstrate that our multi-task approach achieves a mean improvement of 3.2x on held-out test tasks compared to single-task models and consistently outperforms strong behaviour cloning baselines. These results suggest that the development of generalisable MARL agents should prioritise the diversity of the training distribution with varying numbers of agents, providing a roadmap for scaling offline MARL effectively.

补充信息

↑