arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16524cs.LG

反馈归因与表示几何:多智能体强化学习中个体与共享奖励比较的度量

Feedback Attribution and Representation Geometry: Metrics for Comparing Individual and Shared Rewards in MARL

Tasha Pais, Richard Higgins

AI总结:

研究多智能体强化学习中团队平均奖励下,学习表示上奖励归因的特征。提出EffRank/$n$和$D_\text{act}$指标,在SMACv2测试发现观测解释几何特征,奖励归因体现在行为,且指标开销低。

AI中文摘要:

合作多智能体强化学习系统通常使用团队平均奖励,这种反馈归因选择使每个智能体无论个体贡献如何都能获得团队结果。我们研究这是否会在学习表示上留下可测量的特征,无论是几何还是行为方面的。我们提出EffRank/$n$(按智能体数量归一化的有效秩)和$D_\text{act}$(智能体动作分布之间的平均成对KL散度)作为奖励归因效应的低开销诊断指标,并在SMACv2 \texttt{protoss\_5\_vs\_5}中的胜任MAPPO智能体上进行测试,其中单元类型在观测中编码。在观测×奖励归因比较(观测到的单元类型与掩码;个体伤害贡献奖励与共享团队奖励)中,几何特征遵循观测而非奖励。当观测到单元类型时,共享和个体奖励具有相似的EffRank/$n$($0.31{\pm}0.03$对$0.29{\pm}0.02$)和探测准确率($0.75{\pm}0.05$对$0.73{\pm}0.05$,均远高于$1/3$的机会),而$D_\text{act}$在个体奖励下更高($1.23{\pm}0.06$对$1.07{\pm}0.20$)。掩码单元类型会使高于机会的探测信号减少一半以上,在两种奖励方式下均降至$0.49$。简而言之,个体奖励的智能体胜任且可按角色分离,但在SMACv2上,观测解释几何特征,奖励归因主要体现在行为上。因此,几何诊断必须控制观测到的角色信息并测试未直接观测到的持久角色。EffRank/$n$和$D_\text{act}$增加的开销<$5\%$。

英文摘要:

Cooperative multi-agent RL systems routinely use team-averaged rewards, a feedback-attribution choice that gives each agent the team outcome regardless of its individual contribution. We ask whether this leaves a measurable signature, geometric or behavioral, on learned representations. We propose EffRank/$n$ (effective rank normalized by agent count) and $D_\text{act}$ (mean pairwise KL divergence between agents' action distributions) as low-overhead diagnostics for reward-attribution effects, then test them on competent MAPPO agents in SMACv2 \texttt{protoss\_5\_vs\_5}, where unit type is encoded in the observation. In an observation $\times$ reward-attribution comparison (unit type observed vs.\ masked; individual damage-contribution reward vs.\ shared team reward), geometry follows observation rather than reward. With unit type observed, shared and individual rewards have similar EffRank/$n$ ($0.31{\pm}0.03$ vs.\ $0.29{\pm}0.02$) and probe accuracy ($0.75{\pm}0.05$ vs.\ $0.73{\pm}0.05$, both $\gg 1/3$ chance), while $D_\text{act}$ leans higher under individual rewards ($1.23{\pm}0.06$ vs.\ $1.07{\pm}0.20$). Masking unit type cuts the above-chance probe signal by more than half, to $0.49$ in both reward arms. In short: individually rewarded agents are competent and separable by role, but on SMACv2 the observation explains the geometry and reward attribution shows up mainly in behavior. Thus geometric diagnostics must control for observed role information and test persistent roles that are not directly observed. EffRank/$n$ and $D_\text{act}$ add $<$5\% overhead.

补充信息

↑