AI 中文总结
研究定价算法勾结问题,在恒定探索下,通过Q学习过程预期动态推导边界,表明合作策略在时间平均意义上可占主导,该边界能有力预测ε-贪婪Q学习下非背叛主导行为。
AI 中文摘要
定价算法之间的算法勾结引发了对持续超竞争价格及其对社会福利影响的担忧。现有工作主要关注强化学习算法收敛到合作策略的概率,通常假设随着时间推移探索会消失。鉴于实际部署的算法可能会持续探索以适应不断变化的环境,我们研究了恒定探索下的学习动态。在此设置中,相关问题不再是算法是否收敛到特定策略配置,而是算法花费多少时间采用合作策略。即使在具有单周期记忆的重复囚徒困境的基准情况下,这也会产生高维随机学习动态,完整的解析处理难以解决。我们表明合作策略在这种时间平均意义上可以占主导地位,并基于Q学习过程的预期动态推导出预测这种主导地位何时出现的边界。广泛的模拟表明,该边界是ε-贪婪Q学习下非背叛主导行为的有力预测指标。
英文摘要
Algorithmic collusion among pricing algorithms has raised concerns about sustained supra-competitive prices and their implications for social welfare. Existing work has largely focused on the probability that reinforcement-learning algorithms converge to cooperative strategies, typically under the assumption that exploration vanishes over time. Motivated by the observation that algorithms deployed in practice are likely to continue exploring in order to remain adaptive to changing environments, we study learning dynamics under constant exploration. In this setting, the relevant question is no longer whether an algorithm converges to a particular strategy profile, but rather what fraction of time the algorithms spend playing cooperative strategies. Even in the benchmark case of the repeated Prisoner's Dilemma with one-period memory, this yields high-dimensional stochastic learning dynamics, for which a complete analytic treatment is intractable. We show that cooperative strategies can be dominant in this time-averaged sense and derive a boundary predicting when such dominance arises, based on the expected dynamics of the Q-learning process. Extensive simulations show that this boundary is a strong predictor for non-defection-dominated behaviour under epsilon-greedy Q-learning.
Comments35 pages, 14 figures