多智能体夺旗中的观测空间博弈
Games Over Observation Spaces in Multi-Agent Capture the Flag
- Michigan State University(密歇根州立大学)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究多智能体夺旗中,防守方通过限制智能体观测范围以应对固定策略局限,形式化为零和博弈,并提出双预言机算法求解近似均衡,实验验证观测操纵能提升防守性能。
AI中文摘要:
我们考虑在图基环境中的多智能体夺旗(Capture the Flag, CtF)场景,其中一支攻击者队伍试图到达指定的旗帜顶点,而防守队伍则试图拦截他们。在我们的设定中,两支队伍均采用分散式启发式策略运行。虽然攻击队伍可以从多样化的策略库中选择其启发式策略,但防守方仅限于使用单一固定策略。为克服这一限制,一个集中式防守预言机策略性地限制其每个智能体可见的图的部分,以从其固定策略中引发更广泛的涌现行为。我们将这种交互形式化为一个双人零和博弈,其中攻击者在其启发式策略库上进行推理,而防守方则在可见性配置的组合空间上进行推理。为解决这个规模庞大到难以处理的博弈,我们提出了一种双预言机(Double Oracle)算法来寻找近似经验均衡,并通过实验验证了观测操纵可以提升防守方的性能。
英文摘要:
We consider a multi-agent Capture the Flag (CtF) scenario in a graph-based environment, where a team of attackers seeks to reach designated flag vertices while a defending team attempts to intercept them. In our setting, both teams operate using decentralized heuristic policies. While the attacking team may choose its heuristic from a diverse library of policies, the defense is restricted to playing a single fixed policy. To overcome this limitation, a centralized defense oracle strategically restricts the portion of the graph visible to each of its agents in order to elicit a wider range of emergent behaviors from its fixed policy. We formalize this interaction as a two-player zero-sum game, where the attacker reasons over its library of heuristics and the defense reasons over the combinatorial space of visibility profiles. To solve this intractably large game, we propose a Double Oracle algorithm to find approximate empirical equilibria, and we empirically validate that observation manipulation can improve the defense's performance.