arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14559cs.AIcs.LG

何时通信:多智能体强化学习中基于信念分布与KL散度的原则性门控机制

When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL

Teoman Kaman

AI总结:

本文提出一种多智能体强化学习中基于KL散度的原则性通信门控机制,在Predator-Prey和MPE基准测试中,该方法在复杂场景下的性能优于IC3Net等基线方法,还能改进潜在表示以提升智能体协调效果。

AI中文摘要:

多智能体强化学习中的有效通信要求智能体不仅要决定通信的内容,还要决定通信的时机。现有方法要么在每个时间步都进行通信,要么通过REINFORCE策略梯度学习二元门控机制,而REINFORCE是一种高方差信号,会产生不稳定且难以解释的门控行为。本文提出一种原则性替代方案:智能体仅在其学习到的信念分布之间的KL散度超过固定阈值时才进行通信。每个智能体维护一个关于潜在世界状态的信念分布,该分布通过对其LSTM隐藏状态进行softmax计算得到,仅当信念分歧足够大到值得进行信息交换时才会通信。本文在IC3Net的Predator-Prey基准测试上评估该方法,涉及两种环境规模,每种规模使用5个随机种子;同时在MPE simple_spread基准测试上进行评估,对比方法包括IC3Net、CommNet和独立控制器。在10×10的Predator-Prey环境中,IC3Net在所有阈值下的性能均优于KL信念方法;在更困难的20×20 Predator-Prey环境中,对ε∈{0.1,0.3,0.5,1.0}的阈值消融实验显示出倒U型关系:ε=0.5时平均步数为73.84,成功率为42%,而IC3Net的平均步数为75.31,成功率为31%,两者相差1.47步和11个百分点,且ε=0.5时种子方差更紧凑。在MPE环境中,即使门控机制未激活,信念头仍使平均奖励提高12个点,方差降低26倍,这表明该方法有两个正交贡献:信念可收敛时的原则性门控,以及无论门控状态如何都能提升协调效果的改进潜在表示。

英文摘要:

Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I propose a principled alternative: agents communicate only when the KL divergence between their learned belief distributions exceeds a fixed threshold. Each agent maintains a belief distribution over a latent world state computed as a softmax over its LSTM hidden state, and communicates only when belief disagreement is large enough to justify information exchange. I evaluate this approach on the Predator-Prey benchmark from IC3Net \cite{singh2019} across two environment sizes with 5 seeds each, and on MPE simple\_spread \cite{lowe2017}, comparing against IC3Net, CommNet, and an independent controller. On PP 10$\times$10, IC3Net outperforms KL-belief at all thresholds. On the harder PP 20$\times$20, a threshold ablation over $\varepsilon \in \{0.1, 0.3, 0.5, 1.0\}$ reveals an inverted U-shape: $\varepsilon=0.5$ achieves 73.84 average steps and 42\% success rate versus IC3Net's 75.31 steps and 31\%, a gap of 1.47 steps and 11 percentage points with tighter seed variance. On MPE, the belief head improves mean reward by 12 points and reduces variance by 26$\times$ even when gating is inactive, suggesting two orthogonal contributions: principled gating when beliefs can converge, and improved latent representations that benefit coordination regardless.

补充信息

↑