arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

优势聚合而非比率聚合:合作多智能体策略优化的规范形式分析

Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

Zijian Zhao, Sen Li

arXiv 2607.17924首次发表:更新:

发表机构

The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究合作多智能体策略优化中聚合相邻智能体的问题,将设计选择形式化为支持矩阵,证明规范结构,得出聚合应在优势中按耦合邻域大小进行且保持每个智能体比率的设计原则。

AI 中文摘要

以基于近端策略优化(PPO)的方法为例的多智能体策略优化,是合作多智能体强化学习(MARL)的关键分支。一个核心设计问题是聚合多少相邻智能体以有效利用全局信息进行合作。此决策须从优势(哪些智能体的奖励对信用信号有贡献)和比率(哪些智能体的似然比形成裁剪后的重要性权重)两个维度做出。现有方法在这两个轴上占据分散且未充分探索的点。我们将这两个设计选择形式化为支持矩阵$\SA$和$\SR$,并证明了一个规范结构:预期的多智能体策略优化目标仅通过它们的矩阵乘积$\tS=\SR\SA$依赖于对$(\SA,\SR)$。这产生了两个关键结果:(i)冗余性:两个支持矩阵在信号方面是可互换的,意味着没有一种聚合模式本质上更优越。(ii)方差排序:优势以和的形式聚合奖励(在耦合邻域处具有内部偏差 - 方差最优的加性方差),而比率以乘积的形式聚合似然比(随着支持大小呈指数增长的乘性方差,且没有伴随的偏差减少)。由此产生的设计原则很明确:在优势中聚合邻居,其大小与耦合邻域匹配,并保持每个智能体的比率。

英文摘要

Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring agents\footnote{In this paper, "neighbors" refer not only to physical proximity but also to agents whose actions influence one another.} to aggregate in order to effectively utilize global information for cooperation. This decision must be made along two dimensions: in the advantage (which agents' rewards contribute to the credit signal) and in the ratio (which agents' likelihood ratios form the clipped importance weight). Existing methods occupy scattered, underexplored points on these two axes: IPPO treats both separately; MAPPO pairs a team-level advantage with per-agent ratios; HAPPO employs sequential ratios with per-agent advantages; and single-agent reductions operating on factorized joint policies aggregate both into fully joint products. We formalize these two design choices as support matrices $\SA$ and $\SR$, and prove a canonical structure: the expected multi-agent policy optimization objective depends on the pair $(\SA,\SR)$ only through their matrix product $\tS=\SR\SA$. This yields two key consequences: (i) Redundancy: the two support matrices are interchangeable with respect to the signal, meaning neither aggregation pattern is inherently superior.(ii) Variance Ordering: the advantage aggregates rewards as a sum (additive variance with an interior bias-variance optimum at the coupling neighborhood), whereas the ratio aggregates likelihood ratios as a product (multiplicative variance that grows exponentially with support size, with no accompanying bias reduction). The resulting design principle is unambiguous: aggregate neighbors in the advantage, sized to the coupling neighborhood, and keep the ratio per-agent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑