发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出风险规避多群体平均场博弈范式,在模糊集上优化最坏情况奖励,证明均衡存在性与收缩性,并设计风险规避虚拟博弈方案,数值实验验证收敛性。
AI 中文摘要
平均场博弈及其多群体变体的近期进展使得大规模异构多智能体系统能够通过代表性智能体及其相关的平均场分布进行建模。然而,现有方法并未明确考虑其他群体行为中的不确定性。为此,我们引入了一种新范式:风险规避的多群体平均场博弈,其中每个群体在由其他群体子集的平均场流构成的动态可行模糊集上优化最坏情况下的期望奖励。利用占用测度公式以及集值分析的工具,我们在温和假设下建立了多群体博弈的若干理论性质,包括模糊集的几何性质以及一种新颖的风险规避多群体平均场均衡的存在性。进一步,我们推导了熵正则化下不动点算子的收缩性结果,并表明其可用于学习均衡。最后,我们提出了一种风险规避的虚拟博弈方案,并证明尽管最坏情况目标引入了额外的非线性,可剥削性仍衰减至零。我们报告了若干数值实验以说明收敛性和风险规避行为。
英文摘要
Recent advances in mean-field games and its multi-population variants enable large-scale heterogeneous multi-agent systems to be modeled through representative agents and their associated mean-field distributions. However, existing approaches do not explicitly account for uncertainty in the behavior of other populations. To this end, we introduce a new paradigm: risk-averse multi-population mean-field games, where each population optimizes a worst-case expected reward over dynamically feasible ambiguity sets of mean-field flows of a subset of the other populations. Employing an occupation-measure formulation along with tools from set-valued analysis, we establish, under mild assumptions, several theoretical properties of the multi-population game, including the geometric properties of the ambiguity sets and the existence of a novel risk-averse multi-population mean-field equilibrium. Further, we derive contractivity results of the fixed-point operator under entropy regularization and show that it can be utilized to learn the equilibrium. Finally, we propose a risk-averse fictitious-play scheme and show that exploitability decays to zero, despite the additional nonlinearity introduced by the worst-case objective. We report several numerical experiments to illustrate convergence and risk-averse behavior.
CommentsSubmitted to a conference; comments are welcome