AI 中文总结
本文证明未经修改的后验采样PSRL在任意相关先验下实现极小极大贝叶斯遗憾最优,通过共同经验转移参考和贝尔曼方差论证,获得表格型MDP的$\tilde{O}(\sqrt{SAH^3K})$率及线性混合MDP的$\tilde{O}(d\sqrt{H^3K})$率。
AI 中文摘要
强化学习的后验采样(PSRL)是最简单且最有效的探索方法之一,但一个基本问题仍然悬而未决:未经修改的PSRL是否能在不对先验做结构假设的情况下实现极小极大遗憾?我们回答是肯定的。在任意相关先验下,精确的香草PSRL在贝叶斯遗憾的首阶意义上是极小极大最优的。难点在于后验采样的转移模型与其自身的续值相耦合。我们通过一个共同的经验转移参考来克服这一困难,该参考隔离了由此产生的值失配,并借助基于贝尔曼的方差论证来控制它,而无需额外的首阶状态空间因子。对于有限时域、时间非齐次的表格型MDP,且奖励未知随机的情况下,这产生了在任意奖励与转移联合先验下的极小极大遗憾率$\tilde{O}(\sqrt{SAH^3K})$。同样的证明原理在任意联合参数先验下,为线性混合MDP提供了极小极大遗憾率$\tilde{O}(d\sqrt{H^3K})$。
英文摘要
Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayesian regret under arbitrary correlated priors. The difficulty is that a posterior-sampled transition model is coupled with its own continuation value. We overcome this with a common empirical transition reference that isolates the resulting value mismatch and a Bellman-based variance argument that controls it without an extra leading-order state-space factor. For finite-horizon, time-inhomogeneous tabular MDPs with unknown stochastic rewards, this yields the minimax $\widetilde{O}(\sqrt{SAH^3K})$ regret rate under arbitrary joint priors over rewards and transitions. The same proof principle gives the minimax $\widetilde{O}(d\sqrt{H^3K})$ rate for linear-mixture MDPs under arbitrary joint parameter priors.
Comments44 pages