arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09256cs.MA

基于监督网络的分布式团队协调:收敛性、最优性与弹性

Distributed Team Orchestration via Supervisor Networks: Convergence, Optimality, and Resilience

Juntian Zhu, Guanpu Chen, Tongtian Zhu, Miguel de Carvalho, Zhouwang Yang, Fengxiang He

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对带监督网络的零和势团队博弈,提出分布式团队协调算法(DTOA),证明其收敛性与弹性,经实验验证其性能优于基线方法。

中文摘要 AI 辅助

本文研究带有监督网络的零和势团队博弈,其中智能体依赖监督者提供的信念信息而非准确的共同信念。主要挑战在于,由于监督者的信念估计误差以及拜占庭团队对联合行动的误报,这类信念信息可能不准确。我们提出分布式团队协调算法(DTOA),该算法将团队虚拟博弈与基于监督者的分布式信念学习相结合。我们证明了监督者信念估计的收敛性,并确定所诱导的学习动态在团队纳什间隙(TNG)方面收敛至近似团队纳什均衡(TNE)。在拜占庭环境中,我们考虑误报攻击模型并开发了具备拜占庭弹性的DTOA。我们进一步为拜占庭团队识别提供概率保证,并建立诚实团队纳什间隙的渐近界。数值实验验证了理论发现,将DTOA与基线学习方法进行比较,并评估其在马尔可夫决策过程环境中的性能。

英文摘要

In this paper, we study zero-sum potential team games with a supervisor network, where agents rely on supervisor-provided belief information rather than accurate common beliefs. The main challenge is that such belief information can be inaccurate because of supervisors' belief-estimation errors and the misreporting of joint actions by Byzantine teams. We propose the distributed team-orchestrating algorithm (DTOA), which combines team fictitious play with supervisor-based distributed belief learning. We prove the convergence of supervisors' belief estimates and establish that the induced learning dynamics converge to a near team-Nash equilibrium (TNE) in terms of the team-Nash gap (TNG). In the Byzantine setting, we consider a misreporting attack model and develop a Byzantine-resilient DTOA. We further provide probabilistic guarantees for Byzantine-team identification and establish an asymptotic bound on the honest TNG. Numerical experiments illustrate the theoretical findings, compare DTOA with baseline learning methods, and evaluate its performance in a Markov decision process setting.

↑