arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越超竞争结果:深度强化学习最优执行博弈中的合谋行为

Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games

Christos Spyridon Koulouris, Carlo Campajola

arXiv 2610.00619首次发表:更新:

发表机构

University College London; UZH Blockchain Center(伦敦大学学院; 苏黎世大学区块链中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文在最优执行博弈中识别出深度强化学习智能体习得的惩罚机制,该机制通过加速清算惩罚偏离者,为超竞争结果提供了合谋的行为与经济证据。

AI 中文摘要

在本文中,我们通过识别一种习得的惩罚机制来扩展先前关于最优执行博弈中超竞争结果的发现,该机制阻止偏离行为并提供合谋的行为证据。我们在一个双人、有限时域的Almgren-Chriss清算博弈中研究该机制。独立的近端策略优化智能体能够访问情节内的价格和行动历史,其实现的成本低于纳什基准。我们通过针对平均习得清算计划进行训练来识别一种有利可图的偏离,然后将该偏离的首次交易强加给原始智能体之一。对手通过加速清算来回应。这种回应在每次运行和两个玩家角色中都超过了偏离者的收益,同时相对于在相同偏离下不进行惩罚,惩罚者的平均收益基本保持不变。尽管存在更有利可图、惩罚性更低的清算计划,惩罚者仍对偏离者施加更大的损失,同时保持自身的平均收益。匹配偏离和随后的额外卖出在训练过程中先上升后下降,而最终策略保留了有效的惩罚性回应。我们形式化了两个检验:惩罚是否超过偏离的收益,以及交易行为的变化是否足够大以解释所施加的损失。两个检验对所测试的偏离均成立。这些发现共同提供了支持对习得超竞争结果进行合谋解释的行为和经济证据。

英文摘要

In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑