arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04010cs.GTcs.MA

$Q$ 能玩这个游戏:面向连续动作零和马尔可夫博弈的凸-凹函数逼近在线拟合 $Q$-迭代

$Q$ Can Play That Game: Online Fitted $Q$-Iteration for Continuous-Action Zero-Sum Markov Games with Convex-Concave Function Approximation

Kushagra Gupta, Jingqi Li, Cade Armstrong, Lasse Peters, Ross E. Allen, Ufuk Topcu, David Fridovich-Keil

首次发表
浏览论文内容

中文总结 AI 辅助

针对连续动作零和马尔可夫博弈,提出凸-凹函数逼近的在线拟合Q-迭代方法,首次为非LQ博弈提供有限样本保证。

中文摘要 AI 辅助

零和马尔可夫博弈出现在广泛的序贯决策问题中,例如对抗性学习和针对模型不确定性的规划。然而,先前关于零和马尔可夫博弈中学到的状态-动作值函数($Q$-函数)的有限样本保证的工作,在很大程度上局限于有限动作设置,或局限于智能体具有线性动力学和二次奖励(即线性二次(LQ)博弈)的连续动作博弈。将这些保证扩展到更一般的连续动作零和马尔可夫博弈的一个核心挑战是,相关的贝尔曼算子涉及一个极小极大问题,该问题不一定具有可解的鞍点解。为此,我们首先引入一类在智能体动作上凸-凹的$Q$-函数神经网络函数逼近器,这保证了该极小极大问题具有纯策略鞍点。然后,我们研究了使用该函数类的在线拟合$Q$-迭代变体,并据我们所知,首次为具有连续状态和动作的非LQ零和马尔可夫博弈建立了有限样本保证。我们的代码可在以下网址找到:此https URL。

英文摘要

Zero-sum Markov games arise in a wide variety of sequential decision-making problems such as adversarial learning and planning against modeled uncertainties. However, prior work on finite-sample guarantees on the learned state-action value function ($Q$-function) for zero-sum Markov games is largely restricted to finite-action settings, or to continuous-action games in which agents have linear dynamics and quadratic rewards (i.e. linear-quadratic, or LQ, games). A central challenge in extending such guarantees to more general continuous-action zero-sum Markov games is that the associated Bellman operator involves a minimax problem that need not admit a tractable saddle-point solution. To this end, we first introduce a class of neural network function approximators for the $Q$-function that is convex-concave in the players' actions, which guarantees that this minimax problem admits a pure-strategy saddle point. We then study an online variant of fitted $Q$-iteration employing this function class and establish, to the best of our knowledge, the first finite-sample guarantees for non-LQ zero-sum Markov games with continuous states and actions. Our code can be found at https://github.com/CLeARoboticsLab/QCanPlayThatGame.

发表机构

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • The University of California, Berkeley(加州大学伯克利分校)
  • Massachusetts Institute of Technology (MIT) Lincoln Labs(麻省理工学院林肯实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑