动态社交网络中马尔可夫协同进化意见形成博弈的动力学与收敛性
Dynamics and Convergences for Markov Coevolutionary Opinion Formation Games in Dynamic Social Networks
AI总结:
研究动态社交网络中K-NN马尔可夫博弈的收敛性,整合多智能体强化学习和在线学习技术,分析乐观梯度上升算法在一般和马尔可夫博弈中的收敛情况,得出较弱意义上收敛到近似纳什均衡的结果。
AI中文摘要:
虽然协同进化意见形成博弈的确定性变体,如K近邻(K-NN)博弈,在动态社交网络中有时可通过势函数或局部平滑性论证证明其稳定性,但引入随机性会从根本上改变数学情形。在“K-NN马尔可夫博弈”中,网络拓扑通过时变随机选择过程演化。证明这样一个作为一般和马尔可夫博弈特殊情况的系统是否收敛到均衡是一个非常不明显且具有挑战性的理论问题。多智能体强化学习已被证明可在两人零和马尔可夫博弈和马尔可夫势博弈中导出纳什(极小极大)均衡。近期工作表明乐观动力学在一般和马尔可夫博弈中收敛到相关均衡,但其无政府价格界限未知。因此,我们分析在一般和马尔可夫博弈中玩特定无悔算法以收敛到比相关均衡更严格的集合。我们整合了来自Wei等人的多智能体强化学习的收敛分析技术和Anagnostides等人近期工作中的在线学习技术。具体而言,在(一般和)马尔可夫博弈中,由于乐观梯度上升算法的遗憾会有来自Q值的额外正项,处理这些项需要非平凡的额外工作,即设置适当的学习率范围并推导收敛的迭代次数阈值或有界的无政府价格,这与Anagnostides等人主要技术定理中的假设显著不同。我们通过在一般和马尔可夫博弈中玩乐观梯度上升来分析一种较弱意义上的收敛到近似纳什均衡。
英文摘要:
While deterministic variants of the coevolutionary opinion formation games such as the K-Nearest Neighbor (K-NN) game, e.g., in Bhawalkar et al., in a dynamic social network can sometimes be shown to stabilize using potential functions or localized smoothness arguments, introducing stochasticity fundamentally changes the mathematical landscape. In the "K-NN Markov game", network topologies evolve via a time-varying, randomized selection process. Proving whether such a system, as a special case of general-sum Markov games, converges to an equilibrium is a profoundly non-obvious and challenging theoretical question. Multiagent reinforcement learning has been shown to derive Nash (minimax) equilibria in two-player zero-sum Markov games and Markov potential games (along with some price-of-anarchy types of results). In recent work, optimistic dynamics are shown to converge to correlated equilibria in general-sum Markov games while the price-of-anarchy bounds are unknown. We thus analyze playing specific no-regret algorithms in general-sum Markov games for convergence to a stricter set than correlated equilibria. We integrate the convergence analysis techniques from multi-agent reinforcement learning in works of Wei et al. and online learning in a recent work of Anagnostides et al. Specifically in (general-sum) Markov games, since the regret of the optimistic gradient ascent algorithm would have extra positive terms coming from Q-values, taking care of these terms requires non-trivial extra work setting an appropriate range of our learning rate and deriving the threshold on the number of iterations for convergence or a bounded price of anarchy, significantly different from those in the assumption in a main technical theorem of Anagnostides et al. We analyze a weaker sense of convergences to approximate Nash equilibria by playing optimistic gradient ascents in general-sum Markov games.