重复和谐博弈中贝尔曼最优性方程的对称解
Symmetric solution of the Bellman optimality equation for repeated harmony game
浏览论文内容
中文总结 AI 辅助
本研究求解重复和谐博弈中贝尔曼最优性方程的对称解,发现三类解分别对应All-C、赢留输变策略及非平凡策略,并通过数值实验考察了强化学习智能体的实际学习结果。
中文摘要 AI 辅助
在社会困境博弈中,附加奖励或惩罚已被研究作为促进合作的手段。因此,研究这种附加收益会如何改变博弈的理想情境具有重要意义。在本研究中,我们探讨了重复和谐博弈中贝尔曼最优性方程的对称解。计算表明,存在三类对称解。其中一类对应于平凡的All-C策略,另一类对应于囚徒困境博弈中的赢留输变策略。我们还详细讨论了与最后一个解相对应的策略的非平凡行为。此外,我们通过数值实验研究了智能体在强化学习算法下实际学习到的策略。
英文摘要
In social dilemma games, additional rewards or punishments have been studied as means of promoting cooperation. Therefore, it is important to investigate the ideal situation, in which such an additional payoff would change the game. In this study, we investigated the symmetric solution of the Bellman optimality equation for a repeated harmony game. The calculations showed that three types of symmetric solutions exist. One of them corresponds to the trivial All-C strategy, and another to the Win-stay Lose-shift strategy of the prisoners dilemma game. The nontrivial behavior of the strategy corresponding to the last solution is also discussed in detail. In addition, we numerically investigated which strategy the agents actually learn by the reinforcement learning algorithm.
发表机构
- Kindai University(近畿大学)
- Shiga University(滋贺大学)
机构由 AI 辅助整理,请以论文原文为准。