AI 中文总结
本研究通过引入跨实现跨博弈评估方案,系统评估Other-Play算法对实现细节的鲁棒性,发现标准ZSC评估可作为跨实现评估的合理替代指标。
AI 中文摘要
部署在现实场景中的AI智能体必须具备与未遇见过的人类及其他AI智能体协调的能力。零样本协调(ZSC)算法旨在通过指定高级学习规则,让独立设计的智能体在测试时能够相互协调。对ZSC算法的严格评估仍存在困难:理想情况下,需使用每个提出算法的多个独立实现,以反映不同主体解读和实现同一规范时产生的差异。但实际中,ZSC算法几乎仅使用单个实现在不同随机种子上训练,仅少数工作额外改变神经网络架构,这留下了关于算法对规范歧义及实现细节鲁棒性的疑问。本研究首次对该鲁棒性进行系统评估,引入新评估方案——跨实现跨博弈,改变多智能体强化学习(MARL)算法性能受影响的实现细节,并使用该方案评估流行的ZSC算法Other-Play。研究结果令人鼓舞,表明对于Other-Play,标准ZSC评估实际上是更彻底的跨实现评估的合理替代指标。
英文摘要
AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot coordination (ZSC) algorithms aim to achieve this by specifying high-level learning rules such that independently engineered agents can coordinate with each other at test time. Rigorous evaluation of ZSC algorithms remains difficult: ideally, multiple independent implementations of each proposed algorithm must be used, reflecting the variation that arises when independent parties interpret and implement the same specification. In practice, however, ZSC algorithms have almost exclusively been evaluated using a single implementation trained across different random seeds, with only a handful of works additionally varying the neural network architecture. This leaves open questions about robustness to specification ambiguities and implementation details. In this work, we provide the first systematic evaluation of this robustness. We introduce a new evaluation scheme, cross-implementation cross-play, varying implementation details that prior work has shown to affect the performance of multi-agent reinforcement learning (MARL) algorithms, and we evaluate Other-Play, a popular ZSC algorithm, with this scheme. Our findings are encouraging and suggest that, for Other-Play, the standard ZSC evaluation is, in fact, a reasonable proxy for this more thorough cross-implementation evaluation.