线性二次随机博弈中对手不知情下的学习
Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games
浏览论文内容
中文总结 AI 辅助
该研究针对线性二次随机博弈,在对手不知情的最小信息结构下,提出异步分散学习算法并证明其收敛性,应用于动态古诺竞争,发现有限信息学习降低企业利润,公开市场产出可缓解福利损失。
中文摘要 AI 辅助
随着企业越来越多地部署机器学习进行战略决策,理解算法交互已成为运筹学和经济学的核心问题。本文研究无限期非零和线性二次随机博弈中的学习问题,该博弈采用完全解耦的信息结构,其中参与者要么不知道对手,要么策略性地忽略对手,仅观察共同状态和自身的行动历史。在这种最小信息下,我们分析异步分散学习过程,每个参与者独立运行单智能体ε-贪心迭代最小二乘算法。我们证明,尽管参与者无法识别系统参数,其学习动态几乎必然收敛到完全信息纳什均衡,并刻画了收敛速率。随后将该框架应用于具有粘性价格的动态古诺竞争。数值实验验证了理论结果,表明在低和高价格粘性下,有限信息下的学习均会降低企业利润;当价格粘性较高时,总剩余下降且市场集中度上升。公开汇总市场产出会大幅加快收敛并缓解这些福利损失。
英文摘要
As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. This paper studies learning in infinite-horizon, nonzero-sum linear-quadratic stochastic games under a radically uncoupled information structure, where players are either unaware of opponents or strategically oblivious, observing only a common state and their own action history. Under this minimal information, we analyze an asynchronous decentralized learning process in which each player independently runs a single-agent $ε$-greedy iterated least-squares algorithm. We prove that, despite being unable to identify the system parameters, players' learning dynamics converge almost surely to the complete-information Nash equilibrium and characterize the convergence rate. We then apply the framework to a dynamic Cournot competition with sticky prices. Numerical experiments validate the theoretical results and show that learning under limited information reduces firm profits under both low and high price stickiness, while total surplus declines and market concentration increases when price stickiness is high. Publicly revealing aggregate market output substantially accelerates convergence and mitigates these welfare losses.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Imperial College London(伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。