最优多奖励强化学习
Optimal Multi-Reward Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对未知转移且多奖励的有限时域MDP,提出结合MVP、重放样本与间隙乘法权重更新的算法,实现最小最大样本复杂度并匹配下界,同时保证所有奖励的ε-最优策略。
中文摘要 AI 辅助
我们研究了一个转移未知的有限时域马尔可夫决策过程(MDP),其中包含一组有限的已知奖励函数 $\u007br^1, r^2, \backslashldots, r^M\u007d$。目标仅通过在线回合交互为每个奖励输出一个 $\u03b5$-最优策略。性能通过策略误差 $V_{0}^{*, m} - V_{0}^{\backslashwidehat\u007b\u03c0\u007d^m, m}$ 来衡量,其中 $m\u007cin [M]$ 表示奖励函数,$V_{0}^{*, m}=\u007b\backslashmathbb\u007bE\u007d\u007b_s_1\u007bsim \u03bc\u007d[V_{1}^{*, m}(s_1)]$。在此设置下,我们设计了一个可证明高效的算法,建立了最小最大样本复杂度界 $$ O\backslashleft(\backslashfrac\u007bSAH^3\u007d\u007b\u03b5^2\u007d\backslashlog M \backslashmathrm\u007bpolylog\u007d\backslashleft(\backslashfrac\u007bSAH\backslashlog M\u007d\u007b\backslashmin\backslashleft\u007b\u03b5, 1\backslashright\u007d\u03b4\u007d\backslashright)\backslashright)$$ 个回合,且没有额外的预热成本。这匹配信息论下界,相差因子为 $ \backslashmathrm\u007bpolylog\u007d(SAH\backslashlog M/(\backslashmin\backslashleft\u007b\u03b5, 1\backslashright\u007d\u03b4))$。我们的方法结合了三个技术要素。首先,我们将MVP适应于奖励切换学习以构建乐观价值估计。其次,我们使用新鲜的重放样本对候选策略进行保守评估。第三,基于间隙的乘法权重更新利用这些估计之间的差异调整奖励采样分布,将加权学习进度转化为对所有奖励的同时保证。
英文摘要
We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $ε$-optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error $V_{0}^{*, m} - V_{0}^{\widehatπ^{m}, m}$ where $m\in [M]$ represents the reward function and $V_{0}^{*, m}=\mathbb{E}_{s_1\sim μ}[V_{1}^{*, m}(s_1)]$. Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of $$ O\left(\frac{SAH^3}{ε^2}\log M \mathrm{polylog}\left(\frac{SAH\log M}{\min\left\{ε, 1\right\}δ}\right)\right)$$ episodes, with no additional burn-in cost. This matches the information-theoretic lower bound up to a factor of $ \mathrm{polylog}(SAH\log M/(\min\left\{ε, 1\right\}δ))$. Our method combines three technical ingredients. First, we adapt MVP to reward-switching learning to construct optimistic value estimates. Second, we use fresh replay samples to conservatively evaluate the candidate policies. Third, gap-based multiplicative weights updates adjust the reward-sampling distribution using the differences between these estimates, converting weighted learning progress into simultaneous guarantees for all rewards.
发表机构
- Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。