基于Eluder的随机上下文马尔可夫决策过程遗憾界
Eluder-based Regret for Stochastic Contextual MDPs
- Balavatnick school of Computer Science, Tel Aviv University(特拉维夫大学巴拉瓦尼克计算机科学学院)
- Google Research(谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出E-UC$^3$RL算法,在最小假设下实现随机上下文MDP的速率最优遗憾界,并扩展Eluder维度至一般有界度量。
AI中文摘要:
我们提出了E-UC$^3$RL算法,用于随机上下文马尔可夫决策过程(CMDPs)中的遗憾最小化。该算法在可实现函数类和访问\texit{离线}最小二乘及对数损失回归预言机的最小假设下运行。我们的算法是高效的(假设离线回归预言机高效),并享有遗憾保证$ \tilde{O}(H^3 \sqrt{T |S| |A|d_{\mathrm{E}}(\mathcal{P}) \/log (|\mathcal{F}| |\mathcal{P}|/ δ) )}) $,其中$T$是回合数,$S$是状态空间,$A$是动作空间,$H$是视野长度,$\mathcal{P}$和$\mathcal{F}$是分别用于近似上下文相关动态和奖励的有限函数类,$d_{\mathrm{E}}(\mathcal{P})$是$\mathcal{P}$关于Hellinger距离的Eluder维度。据我们所知,我们的算法是第一个在一般离线函数逼近设置下运行的CMDPs的高效且速率最优的遗憾最小化算法。此外,我们将Eluder维度扩展到一般有界度量,这可能具有独立的意义。
英文摘要:
We present the E-UC$^3$RL algorithm for regret minimization in Stochastic Contextual Markov Decision Processes (CMDPs). The algorithm operates under the minimal assumptions of realizable function class and access to \emph{offline} least squares and log loss regression oracles. Our algorithm is efficient (assuming efficient offline regression oracles) and enjoys a regret guarantee of $ \widetilde{O}(H^3 \sqrt{T |S| |A|d_{\mathrm{E}}(\mathcal{P}) \log (|\mathcal{F}| |\mathcal{P}|/ δ) )}) , $ with $T$ being the number of episodes, $S$ the state space, $A$ the action space, $H$ the horizon, $\mathcal{P}$ and $\mathcal{F}$ are finite function classes used to approximate the context-dependent dynamics and rewards, respectively, and $d_{\mathrm{E}}(\mathcal{P})$ is the Eluder dimension of $\mathcal{P}$ w.r.t the Hellinger distance. To the best of our knowledge, our algorithm is the first efficient and rate-optimal regret minimization algorithm for CMDPs that operates under the general offline function approximation setting. In addition, we extend the Eluder dimension to general bounded metrics which may be of separate interest.