arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25701cs.LGcs.MAcs.SYeess.SY

完全拜占庭鲁棒的多智能体强化学习

Fully Byzantine-Resilient Multi-Agent Reinforcement Learning

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

Haejoon Lee, Dimitra Panagou

AI总结:

针对现有拜占庭鲁棒多智能体强化学习方法收敛性能退化的问题,提出FRAC-MARL方法,利用两跳消息冗余识别可靠消息,在拜占庭边攻击下实现参数几乎必然收敛至无攻击极限点,并给出可多项式验证的拓扑条件,在多机器人编队控制中验证有效性。

AI中文摘要:

我们研究了分布式拜占庭鲁棒的演员-评论家多智能体强化学习(AC-MARL),其中智能体通过局部交互共同学习策略。现有方法仅能保证智能体的参数收敛到无攻击极限点的一个邻域内,导致性能下降。我们提出了完全鲁棒的AC-MARL(FRAC-MARL),这是一种去中心化方法,每个智能体利用两跳消息中的冗余来识别可靠消息。在价值和团队奖励函数的线性参数化以及拜占庭边攻击(即对抗行为仅限于通信层)下,我们证明了在时变通信图上,智能体的参数几乎必然收敛到与无攻击情况相同的极限点。我们为方法的收敛引入了一种新颖的拓扑条件,提出了一种系统性的方法来构造此类网络,并证明该条件可以在多项式时间内验证。最后,我们在协作多机器人编队控制任务上演示了我们的方法。

英文摘要:

We study distributed Byzantine-resilient actor-critic multi-agent reinforcement learning (AC-MARL), where agents collectively learn policies through local interactions. Existing methods guarantee convergence of the agents' parameters only to a neighborhood of the attack-free limit points, resulting in degraded performance. We propose Fully Resilient AC-MARL (FRAC-MARL), a decentralized method in which each agent leverages redundancy in two-hop messages to identify reliable messages. Under linear parameterizations of the value and team-reward functions and Byzantine edge attacks, where adversarial behavior is confined to the communication layer, we prove that agents' parameters converge almost surely to the same limit points as in the attack-free case over time-varying communication graphs. We introduce a novel topological condition for the convergence of our method, present a systematic method to construct such networks, and prove that this condition can be verified in polynomial time. Finally, we demonstrate our method on cooperative multi-robot formation control tasks.

补充信息

↑