截断噪声最优响应算法:面向具有安全保证的博弈论学习
Truncated Noisy Best-Response Algorithms: Toward Game Theoretic Learning with Safety Guarantees
浏览论文内容
中文总结 AI 辅助
针对子模最大化多智能体协调问题,提出截断噪声最优响应算法,利用纳什均衡不稳定性,通过异步随机邻域选择动作,给出性能与安全两类递归类界,并揭示二者水床效应关联。
中文摘要 AI 辅助
我们考虑一种博弈论方法来解决具有子模最大化目标的多智能体协调问题。已知对于此类问题,相应博弈的纳什均衡总是在最优值的50%以内,但实现这一最坏情况界值的均衡并不稳定。为了利用这种不稳定性,我们提出了一族算法,称为截断噪声最优响应(TNBR)算法。这些算法的特征灵活,由智能体异步且随机地从其最优响应收益的邻域中选择动作。我们计算了TNBR算法相关马尔可夫链的递归类的界。我们的界分为两类:第一类,“性能”界确保TNBR算法总是具有一个高价值的递归状态;第二类,“安全”界确保TNBR算法永远不会具有任意差的递归状态。此外,这两类界通过一种类似水床效应的机制相互关联:每个具有较差安全保证的博弈必然具有有利的性能保证。
英文摘要
We consider a game theoretic approach to solve multi-agent coordination problems with submodular maximization objectives. It is known for such problems that the Nash equilibria for the corresponding game are always within 50% of the optimal, but that the equilibria which achieve this worst-case bound are not stable. To exploit this instability, we propose a family of algorithms which we call Truncated Noisy Best-Response (TNBR) Algorithms. These algorithms are flexibly characterized by agents asynchronously and stochastically selecting actions from a neighbourhood of their best response payoffs. We compute bounds on the recurrent classes of TNBR algorithms' associated Markov chains. Our bounds fall into two categories: first, "Performance" bounds ensure that TNBR algorithms always have a high-value recurrent state; second, "Safety" bounds ensure that TNBR algorithms never have arbitrarily-bad recurrent states. Furthermore, these two types of bounds are linked by a waterbed-like effect: every game with a poor Safety guarantee necessarily has a favorable Performance guarantee.