发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对未知零和折扣马尔可夫博弈,提出自适应正则化时序差分学习(ARTD),在赌博反馈下实现末轮迭代的 $\widetilde{\mathcal{O}}(t^{-1/4})$ 对偶间隙收敛,显著改进现有速率。
AI 中文摘要
我们研究了在未知两人零和折扣马尔可夫博弈中,带赌博反馈的末轮迭代收敛问题。玩家沿单一轨迹独立学习,不观察彼此的动作。我们开发了自适应正则化时序差分学习(ARTD),在均匀命中时间假设下,该算法对当前策略实现了 $\widetilde{\mathcal{O}}(t^{-1/4})$ 的对偶间隙界,并且以高概率同时对所有轮次和起始状态成立。这改进了 Cai 等人(2023)在相同反馈模型和命中时间假设下,对任意固定 $\nu>0$ 所达到的 $\widetilde{\mathcal{O}}(t^{-1/(9+\nu)})$ 速率。我们的算法不需要知道命中时间界、时间范围或置信水平。为了在价值估计变化时稳定策略学习,我们将快速时序差分平均与有界价值更新分开。我们根据价值估计的进展调整对数障碍正则化,在整个学习过程中控制策略和价值误差。这些机制共同使得实际执行的策略能够快速收敛,即使玩家从赌博反馈中独立学习。
英文摘要
We study last-iterate convergence in unknown two-player zero-sum discounted Markov games with bandit feedback. The players learn independently along a single trajectory without observing each other's actions. We develop Adaptive Regularized TD Learning (ARTD), which achieves a $\widetilde{\mathcal{O}}(t^{-1/4})$ duality gap bound for the current policies under a uniform hitting time assumption, with high probability simultaneously over all rounds and starting states. This improves the $\widetilde{\mathcal{O}}(t^{-1/(9+ν)})$ rate of Cai et al. (2023), for any fixed $ν>0$, under the same feedback model and hitting time assumption. Our algorithm requires no knowledge of the hitting time bound, the time horizon, or the confidence level. To stabilize policy learning as value estimates change, we separate fast temporal difference averaging from bounded value updates. We adapt log-barrier regularization to the progress of value estimation, controlling both policy and value errors throughout learning. Together, these mechanisms enable fast convergence of the policies actually played, even when the players learn independently from bandit feedback.