AI 中文总结
针对通用多层网络中演化博弈的最优控制问题,通过求解哈密尔顿-雅可比-贝尔曼方程得到节点级动态最优激励的解析解,经数值验证后与现有文献结果对比,为促进种群合作提供了新方案。
AI 中文摘要
促进决策者之间的合作行为对社会系统的长期可持续性具有关键意义。在背叛更有利的情况下,激励措施可以促进合作。以往的研究已在由通用单层网络描述社会环境的结构化种群中确定了最优的分散式激励,此处的最优性指成本最小化。本文针对结构化种群在通用网络(这是一个研究较少的案例)中提供了最优激励策略。进一步地,我们研究这样的种群:通过网络结构,每个玩家与一组邻居互动,同时向另一组邻居贡献意见扩散动态,因此我们涵盖了多层网络的情况。为填补这一空白,我们通过求解哈密尔顿-雅可比-贝尔曼方程,为最优控制问题提供了解决方案,并推导了多层网络中激励分配的解析解。通过实现网络化囚徒困境博弈的动态,我们提供了一个反馈控制回路,该回路能随时间产生最优的激励分配。这种动态激励取决于当前的合作状态、博弈层网络、策略扩散层以及收益矩阵设计。我们发现,最优激励是节点级的,对网络上的每个节点都是唯一的,受其相对位置的影响,而相对位置由两层的邻接矩阵决定。我们还发现,最优解并非仅为奖励或惩罚:一个节点可能获得奖励,而另一节点可能同时受到惩罚,这取决于该节点相对于其他节点的当前合作水平。我们为所研究的案例提供了解析解和数值验证,并将我们的结果与现有文献进行了比较。
英文摘要
Promoting cooperative behaviour amongst decision makers has key implications for the long term sustainability of social systems. Incentives can promote cooperation in situations where defection is more favourable. Previous research has identified optimal decentralised incentives in structured populations where the social environment is described by a generic single-layer networks. Optimality is here intended in the sense of cost minimisation. Here, we provide and optimal incentive strategy for structured populations in general network, a rather unexplored case. Further, we look at populations where, through the network structure, each player interact with one set of neighbours while contributing to an opinion diffusion dynamics from a second set of neighbours. Hence we cover the case of multilayer networks. To fill this gap, we provide a solution to the optimal control problem by solving the Hamilton-Jacobi-Bellman equation and derive an analytic solution for distributing incentives in multilayer networks. By implementing the dynamics of the networked Prisoner's Dilemma Game, we provide a feedback control loop that yields an optimal incentive distribution over time. The dynamic incentive depends on the current state of cooperation, the game-layer network, the strategy-diffusion layer, and the payoff matrix design. We found that the optimal incentive is node-wise, unique to each node on the network, influenced by its relative position according to the adjacency matrices of both layers. We also found that the optimal solution is not exclusively reward or punishment. While one node can receive a reward, the other node may receive punishment at the same time, depending on its current level of cooperation, relative to other nodes. We provide analytic solutions and numerical validations for the cases studied, comparing our results to the existing literature.
Comments17 pages, 5 figures, 2 tables