arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无时间范围依赖的强化学习的渐近最优遗憾值

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du

arXiv 2607.19854首次发表:更新:

发表机构

University of Washington; Hong Kong University of Science and Technology(华盛顿大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究有限时间范围齐次表格马尔可夫决策过程的无时间范围遗憾值最小化,提出新算法,证明遗憾值上界\( \tilde O(\sqrt{SAK}+S^8A^3) \),渐近最优且消除\( \log H \)依赖,改进先前结果。

AI 中文摘要

我们研究具有S个状态、A个动作、时间范围H且每个轨迹总奖励受限于1的有限时间范围齐次表格马尔可夫决策过程的无时间范围遗憾值最小化问题。我们提出一种新算法,并证明了一个遗憾值上界\[ \tilde O(\sqrt{SAK}+S^8A^3) \],失败概率为δ,其中K是情节数量,\( \tilde O(\cdot) \)隐藏了\( \mathsf{poly}\log(S,A,K,1/\delta) \)。因此,遗憾值是无时间范围的且渐近最优,与上下文博弈下限\( \Omega(\sqrt{SAK}) \)在对数因子范围内匹配。这完全消除了Zhang等人(2021年)先前的\( \tilde O(\sqrt{SAK\log H}+S^2A\log H) \)保证中的\( \log H \)依赖,并渐近地大幅改进了Zhang等人(2022年)先前的最佳无时间范围遗憾值\( \tilde O(\sqrt{S^9A^3K}) \)。主要技术难点在于,尽管转移核是时间齐次的,但最优值函数\( \{V_h^*\}_{h = 1}^H \)是时间非齐次的。对所有值函数直接使用联合界通常会产生额外的\( \min\{\log H,S\} \)因子。我们通过(i)利用\( V_h^* \)关于h的单调性和(ii)将值函数非平凡地投影到一个S维网格上来避免这个因子。我们的分析依赖于另外三个要素。首先,我们引入一个时间范围截断论证,实现基于奖励的探索并消除单独的无奖励探索阶段的成本。其次,我们设计了一种削减奖励,它既保持了乐观性又保持了规划所需的单调性。第三,我们证明了一个关于时间齐次马尔可夫决策过程总偏差的新界,它以对S的可调多项式依赖来控制削减奖励中的截断方差项,且不依赖于H。这些工具共同产生了一个渐近最优的无时间范围遗憾值保证。

英文摘要

We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$. We propose a new algorithm and prove a regret upper bound \[\tilde O(\sqrt{SAK}+S^8A^3)\] with failure probability $δ$, where $K$ is the number of episodes and $\tilde O(\cdot)$ hides $\mathsf{poly}\log(S,A,K,1/δ)$. Thus, the regret is $H$-free and asymptotically optimal, matching the contextual-bandit lower bound $Ω(\sqrt{SAK})$ up to logarithmic factors. This completely removes the $\log H$ dependence from the previous $\tilde O(\sqrt{SAK\log H}+S^2A\log H)$ guarantee of Zhang et al. (2021), and drastically improves the prior best horizon-free regret $\tilde O(\sqrt{S^9A^3K})$ of Zhang et al. (2022) asymptotically. The main technical difficulty is that the optimal value functions $\{V_h^*\}_{h=1}^H$ are time-inhomogeneous even though the transition kernel is time-homogeneous. A direct union bound over all value functions typically incurs an additional $\min\{\log H,S\}$ factor. We avoid this factor by (i) exploiting the monotonicity of $V_h^*$ in $h$ and (ii) non-trivially projecting the value functions onto an $S$-dimensional grid. Our analysis relies on three additional ingredients. First, we introduce a horizon-truncation argument that enables reward-based exploration and removes the cost of a separate reward-free exploration phase. Second, we design a cutting bonus that preserves both optimism and the monotonicity needed for planning. Third, we prove a new bound on total deviation for time-homogeneous MDPs, which controls the clipped variance terms in the cutting bonus with adjustable polynomial dependence on $S$ and without any dependence on $H$. Together, these tools yield an asymptotically optimal horizon-free regret guarantee.

Comments78 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑