发表机构
Massachusetts Institute of Technology; Woods Hole Oceanographic Institution (WHOI)(麻省理工学院; 伍兹霍尔海洋研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出通过构造状态相关对称正定度量场,在保持所有均衡不变的前提下,使目标均衡稳定、竞争均衡不稳定,从而在非合作可微博弈中实现均衡选择,并在猎鹿博弈和囚徒困境中验证有效性。
AI 中文摘要
度量条件作用可以在不移动可微博弈均衡点的情况下改变吸引学习动力学的均衡。我们研究如何使用状态相关的对称正定(SPD)度量,通过保持指定均衡具有吸引力而使竞争均衡不稳定,从而在已知均衡中进行选择。基于正定乘法下矩阵稳定性的经典结果,我们刻画了这种不稳定化何时可能实现,并展示了将度量限制为对每个玩家独立作用会限制在微分纳什均衡处可实现的稳定性变化。然后,我们通过平滑插值将选定均衡处的稳定度量与竞争均衡处的不稳定度量相结合,产生一个单一的度量场,该度量场保留博弈的每个均衡,同时赋予所需的局部稳定性类型。我们还推导了在离散时间中实现该方法的步长条件。该方法在连续承诺的猎鹿博弈和熵正则化的重复囚徒困境中得到了验证,在这两种设置中都选择了所需的均衡。
英文摘要
Metric conditioning can change which equilibrium attracts learning dynamics without moving the equilibria of a differentiable game. We study how to use state-dependent symmetric positive-definite (SPD) metrics to select among known equilibria by keeping a designated equilibrium attracting while making a rival unstable. Building on classical results on matrix stability under positive-definite multiplication, we characterize when such destabilization is possible and show how restricting the metric to act independently on each player limits the stability changes that can be achieved at differential Nash equilibria. We then combine a stabilizing metric at the selected equilibrium with a destabilizing metric at the rival through smooth interpolation, yielding a single metric field that preserves every equilibrium of the game while assigning the desired local stability types. We also derive a step-size condition for implementing the method in discrete time. The approach is validated on a continuous-commitment stag hunt and an entropy-regularized iterated prisoner's dilemma, where it selects the desired equilibrium in both settings.
Comments8 pages, 2 figures