arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MATES:通过变换观测使冻结的单智能体策略学习多智能体交互

MATES: Learning Multi-Agent Interactions by Transforming Observations for Frozen Single-Agent Policies

Elie Abboud, Oren Gal

arXiv 2609.26010首次发表:更新:

发表机构

University of Haifa(海法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MATES通过为冻结的单智能体策略学习小型观测适配器,在不更新策略参数的情况下实现多智能体协调,仅优化3.5%-7.3%参数且性能优于从头训练。

AI 中文摘要

多智能体强化学习(MARL)通常从头开始训练分散式策略,要求智能体同时获得个体任务能力和协调能力。然而,许多多智能体问题存在一个兼容的单智能体对应版本,其中底层任务可以在隔离环境中学习。我们提出了多智能体观测变换以适配现有单智能体策略(MATES),这是一种输入侧适配框架,适用于其多智能体观测在保留单独任务信息的同时暴露可单独识别的邻居信息的任务。基于多智能体经验,MATES学习一个小型适配器,将该观测映射为冻结的单智能体策略所期望的格式,从而在不更新单智能体策略本身的情况下,产生适合共享环境的动作。MATES保持预训练策略的内部架构不变,并保留底层MARL算法的目标和更新过程。我们使用在线和离线策略算法在终身路径规划、导航和协作发现任务上评估MATES,涵盖离散和连续的观测与动作空间。在所有评估设置中,MATES仅优化全策略训练参数量的3.5%至7.3%,同时始终优于从头开始的MARL训练。它接近完全微调的性能,总体上与基于演示的基线保持竞争力,并在训练中未遇到的团队规模下保持强劲的任务性能。这些结果证明,在这种观测结构下,可以在不修改编码个体能力的策略的情况下学习有效的多智能体行为。

英文摘要

Multi-agent reinforcement learning (MARL) commonly trains decentralized policies from scratch, requiring agents to acquire individual task competence and coordination simultaneously. Yet many multi-agent problems admit a compatible single-agent counterpart in which the underlying task can be learned in isolation. We introduce Multi-Agent Observation Transformation for Existing Single-Agent Policies (MATES), an input-side adaptation framework for tasks whose multi-agent observations preserve the solo-task information while exposing separately identifiable neighbor information. From multi-agent experience, MATES learns a small adapter that maps this observation into the format expected by a frozen single-agent policy, inducing actions suited to the shared environment without updating the single-agent policy itself. MATES leaves the pretrained policy's internal architecture unchanged and retains the objectives and update procedures of the underlying MARL algorithm. We evaluate MATES using both on- and off-policy algorithms on lifelong pathfinding, navigation, and cooperative discovery, spanning discrete and continuous observation and action spaces. Across all evaluated settings, MATES optimizes only 3.5-7.3% as many parameters as full-policy training while consistently outperforming MARL training from scratch. It approaches the performance of full fine-tuning, remains competitive overall with demonstration-based baselines, and retains strong task performance at team sizes not encountered during training. These results provide evidence that, under this observation structure, effective multi-agent behavior can be learned without modifying the policy that encodes individual competence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑