发表机构
Pittsburgh, PA 15213, USA(匹兹堡)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ConventionPlay是一种基于强化学习的方法,通过与具备不同约定适配能力的伙伴群体训练,使智能体主动探测伙伴能力并引导最优联合策略,在多约定兼容的测试伙伴群体中性能优于现有方法。
AI 中文摘要
临时协作通常要求智能体在合作任务中识别并遵守某些共享约定。现有针对临时协作的强化学习(RL)研究聚焦于训练能适配伙伴所建立约定的智能体,但这些方法未考虑以下可能性:部分伙伴仅遵循单一固定约定,而其他伙伴自身可适配多种约定。本文提出ConventionPlay,一种基于RL的方法,通过与学习到的伙伴群体进行训练,该群体中的伙伴在约定间展现出不同程度的适配能力——部分伙伴遵循单一固定约定,其余伙伴则能适配对应任务的部分可能约定。这类仅支持有限约定子集的伙伴的存在,迫使在该群体上训练的智能体主动探测伙伴的能力,并引导伙伴走向双方都能遵循的最有效联合策略。实验结果表明,在与多种约定兼容的测试伙伴群体中,经ConventionPlay训练的智能体性能优于现有临时协作方法。
英文摘要
Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.
CommentsExtension to arxiv:2604.18123. Under review at ICLR 2027