ABRA:一种无法收敛到低质量纳什均衡的算法
ABRA: An algorithm which cannot converge to low-quality Nash equilibria
- University of Colorado at Colorado Springs(科罗拉多大学科罗拉多斯普林斯分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对次模多智能体协调问题,提出噪声与理性参数控制的近似最佳响应算法ABRA,证明其收敛结果避免低质量纳什均衡,且数值模拟验证性能良好。
AI中文摘要:
我们考虑一种博弈论方法来解决具有次模目标的多智能体协调问题。已知对于此类问题,相应博弈的纳什均衡总是在最优值的50%以内。最近的一项工作进一步表明,达到这种最坏情况界限的均衡是不稳定的。利用这一点,我们设计了一种由噪声参数和理性参数控制的近似最佳响应算法(ABRA)。噪声使ABRA能够逃离不良均衡,而理性参数则平衡噪声引起的目标函数的任何退化。我们证明,对于任何两人博弈,如果ABRA收敛到纳什均衡,其系统目标值严格大于最优值的50%加上由噪声参数控制的一项。否则,ABRA收敛到某个循环类:如果一个循环类包含任何产生系统目标低于最优值50%的动作组合,则该类还必须包含最优动作组合或产生系统目标严格大于最优值50%且超出量由噪声参数控制的因子决定的动作组合。ABRA在此类动作组合上花费的时间可以通过理性参数来控制。通过数值模拟,我们表明最小期望目标函数通常远高于最优值的一半。
英文摘要:
We consider a game theoretic approach to solve multi-agent coordination problems with submodular objectives. It is known for such problems that the Nash equilibria for the corresponding game are always within 50% of the optimal. A recent work further shows that the equilibria which achieve this worst-case bound are not stable. Leveraging this, we design an Approximate Best Response Algorithm (ABRA) governed by a noise parameter and a rationality parameter. The noise allows ABRA to escape the bad equilibria and the rationality parameter balances any degradation in the objective function caused by the noise. We show for any two-player game that if ABRA converges to a Nash equilibrium, its system objective value is strictly more than 50% of optimal plus a term controlled by the noise parameter. Otherwise, ABRA converges to some recurrent class: if a recurrent class contains any action profile yielding system objective less than 50% of the optimal, the class must also contain either the optimal action profile or an action profile yielding system objective strictly more than 50\% of the optimal by the same amount in addition to a factor controlled by noise parameter. The time that ABRA spends in such action profiles can be controlled using the rationality parameter. Using numerical simulations, we show that the minimum expected objective function is typically well above half of the optimal.