arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34812cs.LGcs.GT

矩阵博弈中带赌博反馈的对数多项式纳什遗憾

Polylogarithmic Nash Regret in Matrix Games with Bandit Feedback

Yuheng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出乐观支付平衡算法,在有限矩阵博弈中实现对数多项式的纳什遗憾,解决了非唯一均衡下的开放问题,证明观察对手动作足以达到该保证。

中文摘要 AI 辅助

我们研究未知有限矩阵博弈中的纳什遗憾最小化问题,其中支付反馈为赌博式,且能观察到对手动作。我们提出了乐观支付平衡(OPB)算法,该算法在任意自适应对手(包括具有非唯一均衡的博弈)下实现了依赖于实例的$\mathcal{O}(\log^2 T)$纳什遗憾。这解决了Maiti等人(2025)提出的开放问题,将其在赌博反馈下的对数多项式保证从$2\times2$博弈扩展到任意有限维度。为处理非唯一均衡,我们构造了一个留有余地以进行局部调整的参考策略。我们按估计精度对独立支付差异进行排序,并根据不确定性缩放这些调整,使学习者能够利用对手的不平衡来抵消估计成本。我们的结果因此表明,在一般有限矩阵博弈中,观察对手动作足以实现对数多项式的纳什遗憾。

英文摘要

We study Nash regret minimization in unknown finite matrix games with bandit payoff feedback and observed opponent actions. We develop Optimistic Payoff Balancing (OPB), which achieves instance-dependent $\mathcal{O}(\log^2 T)$ Nash regret against arbitrary adaptive opponents, including games with nonunique equilibria. This resolves the open problem posed by Maiti et al. (2025), extending their polylogarithmic guarantee under bandit feedback from $2\times2$ games to arbitrary finite dimensions. To handle nonunique equilibria, we construct a reference strategy that leaves room for local adjustments. We order independent payoff differences by estimation accuracy and scale these adjustments by uncertainty, allowing the learner to exploit the opponent's imbalance to offset estimation costs. Our result thus shows that observing opponent actions suffices for polylogarithmic Nash regret in general finite matrix games.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑