发表机构
Institute for Interdisciplinary Information Sciences, Tsinghua University; Engineering Systems and Design Pillar, Singapore University of Technology and Design(清华大学交叉信息研究院; 新加坡科技设计大学工程系统与设计方向)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一步式流模型OMAF,结合Transformer流策略与近似路径得分代理,实现高效在线多智能体协调,在10个任务上回报提升3.4倍、样本效率提升10.5倍。
AI 中文摘要
多智能体强化学习(MARL)提供了一个通过与环境交互来学习协调行为的强大框架。开发MARL策略需要在复杂且多模态的动作分布的富有表现力的建模与高效的训练和执行之间取得平衡。生成式策略,特别是基于扩散的策略,能够忠实地捕捉复杂和多模态的行为,但昂贵的迭代采样阻碍了它们在在线多智能体环境中的可扩展性。我们提出了一种通过一步式流模型(OMAF)的在线MARL框架,该框架将富有表现力的生成式策略与高效的一步式动作生成相结合。OMAF采用基于Transformer的流策略来捕捉复杂的协调行为,而其近似路径得分代理为同步流策略优化提供了一条原则性途径。为了实现稳定且样本高效的学习,我们进一步开发了一种联合优化方案,将软最大值Q值估计与联合流策略目标相结合,以进行协调策略学习。通过消除迭代采样,OMAF在不牺牲策略表现力的情况下大幅降低了训练开销。在来自MPE和MAMuJoCo的10个标准任务上的大量实验表明,OMAF始终实现卓越的性能,与基线方法相比,回报最高提升3.4倍,样本效率提升10.5倍。这些结果验证了OMAF作为一种富有表现力且计算高效的一步式流策略范式在在线MARL中的有效性。
英文摘要
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.