发表机构
Massachusetts Institute of Technology; Aerospace Control Laboratory; Laboratory of Information and Decision Systems(麻省理工学院; 航空控制实验室; 信息与决策系统实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究非对称信息下对抗团队博弈,提出概率鲁棒极小极大遗憾均衡(PR-MRE),结合极小极大遗憾推理的无分布鲁棒性与概率信息,通过鲁棒双线性规划及新型元求解器 PRMRE-PSRO 实现策略学习,实验证明其在隐藏类型上性能更优,行为更稳健。
AI 中文摘要
非对称信息下的对抗团队博弈,如对抗路径寻找、目标搜索和图上可达性博弈,需要对隐藏对手类型和欺骗具有鲁棒性的策略。现有风险中性解概念对分布变化敏感,分布鲁棒方法仅在规定模糊集内提供保证。为解决这些限制,引入概率鲁棒极小极大遗憾均衡(PR-MRE),它结合了极小极大遗憾推理的无分布鲁棒性和名义类型分布的概率信息。PR-MRE在类型空间的高置信子集中最小化最坏情况遗憾,防止概率质量的战略重新分配,避免完全无分布方法的保守性。对于正规形式贝叶斯博弈,PR-MRE可表述为鲁棒双线性规划并导出可处理的半定松弛。将此松弛应用于鲁棒双 oracle 框架内的新型元求解器 PRMRE-PSRO,通过深度强化学习最佳响应实现基于群体的近似 PR-MRE 策略学习。图结构对抗团队博弈实验表明,与风险中性均衡解相比,PR-MRE 在隐藏类型上发现了具有显著改进的最坏情况性能的策略,在战略分布变化下行为更稳健。
英文摘要
Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs, require strategies that are robust to hidden opponent types, such as a hidden goal flag, and to deception. Under asymmetric information, deception is seen as strategic shifts in the type distribution such that the omniscient opponent can collude with Nature and condition its play on the observed type. Existing risk-neutral solution concepts, such as Bayesian Nash equilibrium (BNE), are sensitive to distribution shifts, while distributionally robust approaches provide guarantees only within a prescribed ambiguity set. To address these limitations, we introduce Probabilistically Robust Minimax-Regret Equilibrium (PR-MRE), a novel equilibrium concept that combines the distribution-free robustness of minimax-regret reasoning with probabilistic information from a nominal type distribution. PR-MRE minimizes worst-case regret over a high-confidence subset of the type space, providing protection against strategic redistribution of probability mass while avoiding the conservatism of fully distribution-free approaches. We show that, for normal-form Bayesian games, PR-MRE can be formulated as a robust bilinear program and derive a tractable semidefinite relaxation. We then adapt this relaxation into a novel meta-solver within a robust double-oracle framework, PRMRE-PSRO, enabling population-based learning of approximate PR-MRE strategies via deep reinforcement learning best responses. Experiments on graph-structured adversarial team games demonstrate that PR-MRE discovers strategies with substantially improved worst-case performance across hidden types compared to risk-neutral equilibrium solutions, resulting in more robust behavior under strategic distribution shifts.
Comments29 pages, 11 figures, 6 tables. Submitted to Transactions on Machine Learning Research (TMLR)