发表机构
ETH Zurich; Imperial College London(苏黎世联邦理工学院; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对不完全均衡观测下的逆向博弈学习问题,提出一种凸的博弈论次优损失,并开发镜像下降算法,在噪声数据下保持准确性,优于逆变分不等式方法。
AI 中文摘要
许多现代系统涉及多个智能体的策略性互动。在此类环境中,观察到的行为通常反映了在部分已知效用下的均衡行为。从数据中恢复这些隐藏效用——逆向博弈论的核心目标——对于预测、反事实分析和机制设计至关重要。然而,现有的基于逆变分不等式的方法对噪声和不一致的均衡观测高度敏感,从而限制了其适用性。在本文中,我们通过引入一种博弈论次优损失来解决这一问题,该损失衡量了智能体通过单方面偏离观察到的策略剖面所能获得的总体效用增益。首先,我们证明该损失是凸的,并且可以高效地分解为每个智能体的最优反应。其次,我们证明该损失介于可预测性损失和逆变分不等式损失之间,使其成为均衡预测的易处理替代。第三,我们开发了一种镜像下降算法来最小化该损失,并在异构网络化古诺竞争中证明,在噪声观测和不一致均衡数据下,我们的方法仍保持准确性,而逆变分不等式方法则产生退化估计。
英文摘要
Many modern systems involve the strategic interaction of multiple agents. In such settings, observed actions typically reflect equilibrium behavior under utilities that are only partially known. Recovering these hidden utilities from data - the central goal of inverse game theory - is key for prediction, counterfactual analysis, and mechanism design. However, existing approaches based on inverse variational inequalities are highly sensitive to noisy and inconsistent equilibrium observations, thus limiting their applicability. In this paper, we resolve this issue by introducing a game-theoretic suboptimality loss that measures the aggregate utility gain players could obtain by unilaterally deviating from an observed strategy profile. First, we show that this loss is convex and admits an efficient decomposition into player-wise best-responses. Second, we show this loss is sandwiched between the predictability loss and the inverse variational inequality loss, making it a tractable surrogate for equilibrium prediction. Third, we develop a mirror descent algorithm to minimize it and demonstrate on a heterogeneous networked Cournot competition that our approach remains accurate under noisy observations and inconsistent equilibrium data while inverse variational inequality methods produce degenerate estimates.