AI 中文总结
该研究针对μ-重置交互协议下的策略学习,明确了策略可实现性对样本复杂度的作用,得出不同集中性条件下时间范围H对应的样本复杂度上下界。
AI 中文摘要
我们研究了Kakade和Langford[KL02]提出的μ-重置交互协议下基于策略的强化学习。该交互协议使学习者除了初始分布外,还能从给定的探索性重置分布μ中采样轨迹。我们解决了[KLS25]提出的关于策略可实现性在该问题样本复杂度中的作用的问题。关键的是,对时间范围H的依赖由所假设的重置分布的覆盖概念决定。在有界全策略集中性条件下,我们证明了指数级的exp(Ω(H))样本复杂度下界;在有界推前集中性条件下,我们证明了对时间范围的依赖被紧密刻画为exp(Θ(√H))。
英文摘要
We study policy-based reinforcement learning under the $μ$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $μ$, in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon $H$ is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a $\exp(Ω(H))$ sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as $\exp(Θ(\sqrt H))$.
Commentscomments welcome