发表机构
Bilkent University(比尔肯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种去中心化最优均衡学习算法,通过交换时间戳堆叠表而非原始数据,在动态通信网络上实现社会最优均衡选择,并给出有限时间对数遗憾保证。
AI 中文摘要
本文研究了在动态通信网络上,有限正规形式博弈中社会最优均衡的去中心化学习问题。每个智能体仅观察自身实现的收益,事先不知道博弈模型,且只能通过低带宽消息与随时间变化的邻居通信。我们提出了网络化去中心化最优均衡学习动力学,其中智能体从局部收益比较中生成随机化的语义内容/不满信号,并交换带时间戳的堆叠表,而非原始动作、收益信息或局部估计/参数。该方法将表融合与时间多数重构相结合,以缓解动态通信的影响,同时保持完全去中心化的运行。我们建立了有限时间对数遗憾界,并采用同相探索扰动,在功利主义与比例公平社会福利目标下实现最优均衡选择。仿真结果进一步表明,所提方法能在动态通信网络上有效选择社会合意的均衡。
英文摘要
This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks. Each agent observes only its own realized payoffs, does not know the game a priori, and can communicate only with time-varying neighbors using low-bandwidth messages. We propose networked decentralized optimal equilibrium learning dynamics in which agents generate randomized semantic content/discontent signals from local payoff comparisons and exchange time-stamped time-stacked tables rather than raw actions, payoff information or local estimates/parameters. The method combines table fusion with temporal majority reconstruction to mitigate dynamic communication while preserving fully decentralized operation. We establish finite-time logarithmic regret guarantees, with an in-phase exploration perturbation, for optimal equilibrium selection under utilitarian and proportional-fair social welfare objectives. Simulation results further show that the proposed approach can effectively select socially desirable equilibria over dynamic communication networks.
CommentsExtended version of the paper: S. T. Kiremitci and M. O. Sayin, "Decentralized optimal equilibrium learning over dynamic networks", to appear in the Proceedings of the IEEE Conference on Decision and Control, 2026