利用在线函数逼近实现对抗性上下文 MDP 的高效速率最优遗憾
Efficient Rate Optimal Regret for Adversarial Contextual MDPs Using Online Function Approximation
- Tel Aviv University(特拉维夫大学)
- Google Research(谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出 OMG-CMDP! 算法,在在线回归预言机与可实现函数类假设下,实现对抗性上下文 MDP 的高效、鲁棒且速率最优遗憾保证。
AI中文摘要:
我们提出了用于对抗性 Contextual MDP(上下文马尔可夫决策过程)中遗憾最小化的 OMG-CMDP! 算法。该算法在可实现函数类以及可访问在线最小二乘和对数损失回归预言机的最小假设下运行。我们的算法是高效的(假设在线回归预言机高效)、简单的,并且对逼近误差具有鲁棒性。它享有 $\widetilde{O}(H^{2.5} \sqrt{ T|S||A| ( \mathcal{R}(\mathcal{O}) + H \log(δ^{-1}) )})$ 的遗憾保证,其中 $T$ 是回合数,$S$ 是状态空间,$A$ 是动作空间,$H$ 是视野,而 $\mathcal{R}(\mathcal{O}) = \mathcal{R}(\mathcal{O}_{\mathrm{sq}}^\mathcal{F}) + \mathcal{R}(\mathcal{O}_{\mathrm{log}}^\mathcal{P})$ 是回归预言机遗憾之和,分别用于逼近依赖上下文的奖励和动态。据我们所知,我们的算法是首个在在线函数逼近这一最小标准假设下运行的、用于对抗性 CMDP 的高效且速率最优的遗憾最小化算法。
英文摘要:
We present the OMG-CMDP! algorithm for regret minimization in adversarial Contextual MDPs. The algorithm operates under the minimal assumptions of realizable function class and access to online least squares and log loss regression oracles. Our algorithm is efficient (assuming efficient online regression oracles), simple and robust to approximation errors. It enjoys an $\widetilde{O}(H^{2.5} \sqrt{ T|S||A| ( \mathcal{R}(\mathcal{O}) + H \log(δ^{-1}) )})$ regret guarantee, with $T$ being the number of episodes, $S$ the state space, $A$ the action space, $H$ the horizon and $\mathcal{R}(\mathcal{O}) = \mathcal{R}(\mathcal{O}_{\mathrm{sq}}^\mathcal{F}) + \mathcal{R}(\mathcal{O}_{\mathrm{log}}^\mathcal{P})$ is the sum of the regression oracles' regret, used to approximate the context-dependent rewards and dynamics, respectively. To the best of our knowledge, our algorithm is the first efficient rate optimal regret minimization algorithm for adversarial CMDPs that operates under the minimal standard assumption of online function approximation.