arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

排列鲁棒性并不足够:多智能体Transformer策略中的动作坍缩

Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies

Amit Thakur, Mukesh Singhal

arXiv 2610.02848首次发表:更新:

发表机构

University of California, Merced(加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示多智能体Transformer策略在排列鲁棒性评估中的陷阱,提出结合动作坍缩诊断,并发现弱等变性正则化可在保持动作多样性的同时提升鲁棒性。

AI 中文摘要

Transformer策略因其自注意力机制能够建模智能体之间的交互,在多智能体机器人学习中颇具吸引力。然而,多智能体团队是无序的,而Transformer通常将智能体作为有序的令牌序列进行处理。我们研究了这种不匹配如何在智能体顺序排列下影响协作导航策略。我们的结果表明,仅凭较低的排列误差可能具有误导性:策略看似鲁棒,可能仅仅是因为所有智能体选择了相同的动作。因此,我们使用排列一致性度量和动作坍缩诊断(包括动作多样性、相同动作比例和最大动作频率)来评估策略。PPO-ID基线产生非坍缩行为,但仍对顺序敏感,而强等变性正则化仍可能诱导同质行为。弱等变性惩罚在N=3个智能体的团队中提高了鲁棒性,同时保留了更多样化的动作,而N=4个智能体的团队则需要显著更小的正则化权重。这些发现表明,多智能体Transformer策略不仅应通过回报和排列鲁棒性来评估,还应检查其是否维持非坍缩、差异化的多智能体行为。

英文摘要

Transformer policies are attractive for multi-agent robot learning because self-attention can model interactions among agents. However, multi-agent teams are unordered, while transformers typically process agents as ordered token sequences. We study how this mismatch affects cooperative navigation policies under agent-order permutations. Our results show that low permutation error alone can be misleading: policies may appear robust simply because all agents choose the same action. We therefore evaluate policies using both permutation-consistency metrics and action-collapse diagnostics, including action diversity, same-action fraction, and maximum action frequency. A PPO-ID baseline yields non-collapsed behavior but remains order-sensitive, while strong equivariance regularization can still induce homogeneous behavior. A weak equivariance penalty improves the robustness while preserving more diverse actions for teams with \(N=3\) agents, whereas teams with \(N=4\) agents require substantially smaller regularization weights. These findings suggest that multi-agent transformer policies should be evaluated not only by return and permutation robustness, but also by whether they maintain non-collapsed, differentiated multi-agent behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑