发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对未知零和矩阵博弈,提出一种高效算法,实现每轮高概率的极小极大最后迭代收敛,改善维度依赖并匹配下界。
AI 中文摘要
我们研究了在具有赌博机支付反馈和观测对手动作的未知两人零和矩阵博弈中的最后迭代收敛问题。对于每个玩家有$d$个动作的博弈,我们开发了一种算法,在每一轮$t$同时以高概率实现$\tilde{\mathcal{O}}(\sqrt{d/t})$的对偶间隙。这比先前已知最佳保证的维度依赖改善了$d^{3/2}$倍。该速率匹配标准的赌博机下界,在动作数量和轮数方面(对数因子内)确立了极小极大最优性。该算法计算高效,每轮仅需$\mathcal{O}(d)$的时间和内存。我们的技术贡献是自适应平均和修正指数权重的联合设计,以吸收估计方差,以及一个限制阶段持续时间的势函数论证。
英文摘要
We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/t})$ with high probability, simultaneously at every round $t$. This improves the dimension dependence of the best previously known guarantee by a factor of $d^{3/2}$. The rate matches a standard bandit lower bound, establishing minimax optimality in both the number of actions and the number of rounds, up to logarithmic factors. The algorithm is computationally efficient, requiring only $\mathcal{O}(d)$ time and memory per round. Our technical contribution is a joint design of adaptive averaging and corrected exponential weights that absorbs estimation variance, together with a potential argument that bounds phase durations.