缓解重复博弈中的报复性算法合谋
Mitigating Retaliatory Algorithmic Collusion in Repeated Games
浏览论文内容
中文总结 AI 辅助
针对重复博弈中强化学习智能体自发形成的报复性合谋,本文提出CURB框架,通过惩罚策略间的总变差距离信号,将合谋均衡转化为平凡均衡,并在多种博弈中显著减少合谋。
中文摘要 AI 辅助
在重复交互中,被训练以最大化自身奖励的强化学习智能体,可以在没有通信或共享设计的情况下,收敛到类似于显式合谋的超竞争结果。现有的缓解方法在很大程度上局限于特定的经济环境,如双边平台和拍卖,这留下了如何为一般重复博弈设计干预措施的问题。我们通过将先前关于Q学习合谋工作的经验观察与简单惩罚码(SPCs)的经典理论之间的关联形式化,来解决这一空白。我们证明任何非平凡的SPC都会在智能体的策略中诱发可量化的条件依赖性,这种依赖性可通过智能体在合作与背叛历史中动作分布之间的总变差距离来检测。基于这一联系,我们提出了CURB(通过奖励塑形和信念注入来解除合谋),一种奖励塑形框架,它在Q学习期间惩罚这种总变差(TV)距离信号,并保证将动力学的任何SPC不动点转化为平凡不动点,从而排除由惩罚威胁维持的合谋均衡。实验上,CURB在Bertrand和Cournot竞争重复博弈中显著减少了Q学习智能体的合谋。我们进一步证明CURB可扩展到Bertrand竞争中的深度Q网络智能体,表明该机制可推广到表格Q学习之外。
英文摘要
Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces a quantifiable conditional dependence in agents' policies, detectable via the total variation distance between an agent's action distributions across cooperation and defection histories. Building on this connection, we propose CURB (Collusion Unwinding via Reward shaping and Belief injection), a reward-shaping framework that penalizes this Total Variation (TV) distance signal during Q-learning and is guaranteed to convert any SPC fixed point of the dynamics into a trivial one, thus precluding collusive equilibria sustained by punishment threats. Empirically, CURB substantially reduces collusion by Q-learning agents in both Bertrand and Cournot Competition Repeated Games. We further demonstrate that CURB extends to deep Q-network agents in Bertrand competition, suggesting the mechanism generalizes beyond tabular Q-learning.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。