arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31281cs.AI

MA-WAM:用于测试时规划的多智能体世界动作模型

MA-WAM: Multi-Agent World-Action Model for Test-Time Planning

Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen, Guojie Wang, Hejun Wu

中文总结 AI 辅助

针对多智能体协作中联合动作依赖建模难的问题,提出MA-WAM测试时规划框架,利用世界模型预测联合动作后果,在30个离线MARL任务上平均相对增益达22.0%以上,且开销极小。

中文摘要 AI 辅助

多智能体协作任务要求不同智能体同时执行联合动作,且每个智能体的动作既影响其他智能体的观测,也影响其响应。因此,需要一个世界模型来预测所有智能体联合动作所产生的团队回报。一种朴素的扩展方式是,在逐步预测团队回报时,将单智能体世界模型直接应用于每个智能体的动作。然而,这种扩展未能捕捉多智能体同时动作之间的依赖关系。我们提出了多智能体世界动作模型(MA-WAM),这是一种测试时规划框架,能够使冻结的多智能体流策略评估候选联合动作的未来结果。据我们所知,MA-WAM是首个用于多智能体流策略的测试时世界模型规划器。MA-WAM根据跨智能体依赖关系预测每个联合动作的后果,并实现高效的候选评分。在MAMuJoCo、SMAC和MPE上的30个离线多智能体强化学习(MARL)设置中,MA-WAM相比直接执行实现了22.0%的平均相对增益,相比均匀动作选择实现了25.6%的平均相对增益。在A100 GPU上的标准评估协议下,MA-WAM增加了12.1毫秒的延迟,占实测生成与评分时间的2.5%。

英文摘要

Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an extension fails to capture the dependencies among the simultaneous actions of multiple agents. We propose Multi-Agent World-Action Model (MA-WAM), a test-time planning framework that enables a frozen multi-agent flow policy to evaluate futures of candidate joint actions. To our knowledge, MA-WAM is the first test-time world-model planner for multi-agent flow policies. MA-WAM predicts the consequences of each joint action according to cross-agent dependencies and enables efficient candidate scoring. Across 30 offline multi-agent reinforcement learning (MARL) settings on MAMuJoCo, SMAC, and MPE, MA-WAM achieves mean relative gains of 22.0% over direct execution and 25.6% over uniform action selection. Under the standard evaluation protocol on an A100 GPU, MA-WAM adds 12.1 ms, accounting for 2.5% of the measured generation-and-scoring time.

补充信息

↑