arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

博弈论均衡与无政府价格的悖论

Paradoxes of Game Theoretic Equilibria and Price of Anarchy

Georgios Piliouras, Ian Gemp, Siqi Liu, Luke Marris

arXiv 2607.11752首次发表:更新:

发表机构

Google Deepmind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究博弈论均衡与无政府价格的悖论,指出将多智能体学习简化为静态均衡和黑箱后悔分析存在问题,通过分析内部纳什均衡、拥堵博弈等情况揭示相关不足,并研究非原子极限下的均衡变化,强调需重新评估基于动态的最坏情况均衡框架。

AI 中文摘要

几十年来,静态解概念(纳什、相关和粗相关均衡)以及无政府价格(PoA)构成了算法博弈论的基石,无悔学习被证明能快速收敛到此类博弈论均衡。我们表明,将多智能体学习简化为静态均衡和黑箱后悔分析会掩盖潜在的动态不均衡和博弈论界限。首先,内部纳什均衡缺乏\(C^1\)向量场信息,这意味着智能体无法区分一致激励与严格对立激励。继承这种几何结构,决定稳健PoA界限的最坏情况纯纳什均衡表现为拓扑不稳定的严格鞍点,在典型拥堵博弈中,表现为几乎处处由严格劣势策略支撑的全局排斥子。将效率保证锚定到这些不稳定状态会导致代数敏感性;我们证明,考虑所有严格正仿射成本会使PoA无界。此外,将学习轨迹投影到相关博弈的离散单纯形上会系统地容纳不可理性化行为。通过粗相关均衡或近端细化评估动态无法排除严格劣势策略。而且,最优\(O(1/T)\)交换后悔最小化无法排除宏观湍流,即使在最小博弈中也表现为混沌极限集。最后,我们研究拥堵博弈的非原子极限。尽管被认为具有高度稳定性且具有紧密的次线性\(\Theta(p / \ln p)\) PoA界限(其中\(p\)是多项式次数),但我们证明,在离散时间学习下,唯一均衡会退化为李 - 约克混沌和全局吸引子,其时间平均无效率会以\(2^p\)的速度指数下降。这些结果需要重新评估基于动态的度量的最坏情况均衡框架。

英文摘要

For decades, static solution concepts (Nash, Correlated, and Coarse Correlated Equilibria) and the Price of Anarchy (PoA) have formed the bedrock of algorithmic game theory, with no-regret learning proving fast convergence to such game-theoretic equilibria. We show that reducing multi-agent learning to static equilibrium and black-box regret analysis obscures underlying dynamic disequilibrium and game theoretic bounds. First, interior Nash equilibria lack $C^1$ vector field information, meaning agents cannot distinguish aligned from strictly opposing incentives. Inheriting this geometry, the worst-case pure Nash equilibria dictating robust PoA bounds manifest as topologically unstable strict saddles, and in canonical congestion games, as global repellers supported on almost everywhere strictly dominated strategies. Anchoring efficiency guarantees to these unstable states causes algebraic sensitivity; we prove that accommodating all strictly positive affine costs renders the PoA unbounded. Furthermore, projecting learning trajectories onto the discrete simplex of correlated play systematically accommodates non-rationalizable behavior. Evaluating dynamics via Coarse Correlated Equilibria or proximal refinements fails to preclude strictly dominated strategies. Moreover, optimal $O(1/T)$ swap-regret minimization does not preclude macroscopic turbulence, manifesting as chaotic limit sets even in minimal games. Finally, we examine the non-atomic limit of congestion games. Though considered highly stable with tight sub-linear $Θ(p/\ln p)$ PoA bounds (where $p$ is the polynomial degree), we prove that under discrete-time learning, the unique equilibrium destabilizes into Li-Yorke chaos and global attractors whose time-averaged inefficiency degrades exponentially as $2^p$. These results necessitate re-evaluating worst-case equilibrium frameworks for dynamically grounded metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑