arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15258math.OCmath.PR

用于势单调遍历平均场博弈的自虚构博弈方法

Self-fictitious-play for Potential Monotone Ergodic Mean-field Games

Yupeng Bai, Mathieu Laurière, Zhenjie Ren, Songbo Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对环面上的遍历单调势平均场博弈,提出自虚构博弈动力学,证明其压缩性与唯一不变律,该不变律接近纳什均衡,误差为信念更新率平方根,线性二次示例与数值实验验证结果。

中文摘要 AI 辅助

我们研究遍历、势、单调平均场博弈(MFGs)中的长期学习问题,采用自虚构博弈(SFP)动力学,将最优控制扩散过程与缓慢演化的信念耦合。每一时间步,状态遵循与当前信念关联的最优反馈,而信念更新使用玩家自身的经验占据测度而非总体分布。针对环面上的遍历单调势MFGs,我们证明SFP动力学是压缩的,且存在唯一不变律;进一步表明该不变律在数量上接近MFG纳什均衡,误差阶为信念更新率的平方根。证明结合了遍历哈密尔顿-雅可比-贝尔曼方程的一致时间正则性估计与基于Lasry-Lions散度的能量论证。线性二次示例显示该速率是尖锐的,数值实验验证了预测的缩放关系。

英文摘要

We investigate long-time learning in ergodic, potential, monotone mean-field games (MFGs) via a self-fictitious-play (SFP) dynamics coupling an optimally controlled diffusion with a slowly evolving belief. At each time, the state follows the optimal feedback associated with the current belief, while the belief is updated using the player's own empirical occupation measure rather than the population distribution. For ergodic monotone potential MFGs on the torus, we prove that the SFP dynamics is contractive and admits a unique invariant law. Moreover, we show that this invariant law is quantitatively close to the MFG Nash equilibrium, with an error of order equal to the square root of the belief-update rate. The proof combines uniform-in-time regularity estimates for the ergodic Hamilton-Jacobi-Bellman equation with an energy argument based on the Lasry-Lions divergence. The linear-quadratic example shows that this rate is sharp, and the numerical experiments illustrate the predicted scaling.

↑