arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开放团队多智能体强化学习中的换员正交信用分配

Turnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning

Amit Thakur, Mukesh Singhal

arXiv 2610.02847首次发表:更新:

发表机构

University of California, Merced(加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开放团队中换员与动作信用混淆问题,提出TOCA价值分解方法,分离动作、换员及交互效应,通过置换不变评论家与反事实信用信号提升动态合作场景下的学习稳健性与回报。

AI 中文摘要

开放团队多智能体强化学习研究在回合过程中智能体可能加入、离开或被替换的合作系统。在此类设置中,团队回报的变化既源于智能体选择有用动作,也源于活跃群体自身的变化。标准集中式评论家与共享优势函数常将这两种效应混合为单一标量信用信号,导致存活的智能体可能因自身无法控制的外生换员事件而受到奖励或惩罚。我们提出换员正交信用分配(TOCA),一种针对开放团队的价值分解方法,将动作效应、纯换员效应以及动作-换员交互效应分离。在外生换员条件下,事件条件价值允许一种中心化分解,其事件条件基线去除纯换员成分,同时保留使团队对未来替换具有鲁棒性的动作信用。我们通过一个对可变规模智能体集合和事件令牌具有置换不变性的集中式评论家来实现这一思想,并推导出反事实的每智能体信用信号以及用于高方差控制环境的软加权交互变体TOCA-β。受控诊断实验表明,TOCA在回报上优于事件感知的MAPPO风格评论家,且移除交互信用会显著损害性能。在仅替换的Dynamic Spread基准中,TOCA-β在高换员率下取得最佳平均回报,并优于其无交互消融版本。这些结果表明,明确分离换员与动作信用是动态合作团队中稳健学习的有用原则。

英文摘要

Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standard centralized critics and shared advantages often mix these two effects into one scalar credit signal, allowing surviving agents to be rewarded or penalized for exogenous turnover events outside their control. We introduce turnover-orthogonal credit assignment (TOCA), a value decomposition for open teams that separates action effects, pure turnover effects, and action--turnover interactions. Under exogenous turnover, the event-conditioned value admits a centered decomposition whose event-conditioned baseline removes the pure turnover component while preserving credit for actions that make the team robust to future replacements. We instantiate this idea with a permutation-invariant centralized critic over variable-size agent sets and event tokens, and derive both a counterfactual per-agent credit signal and a softly weighted interaction variant, TOCA-$β$, for high-variance control environments. Controlled diagnostic experiments show that TOCA improves return over event-aware MAPPO-style critics and that removing interaction credit substantially hurts performance. In a replacement-only Dynamic Spread benchmark, TOCA-$β$ achieves the best mean return at high turnover rates and improves over its no-interaction ablation. These results suggest that explicitly separating turnover from action credit is a useful principle for robust learning in dynamic cooperative teams.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑