发表机构
Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分析结合单步TD更新的策略镜像下降算法,证明其在有限折扣MDP中全局线性收敛,并在随机设置下给出无需轨迹重置的样本复杂度保证。
AI 中文摘要
策略镜像下降(PMD)在正则化马尔可夫决策过程(MDPs)中享有快速收敛性,但现有保证往往依赖于精确或日益精确的策略评估。我们分析了与由单步时间差分(TD)更新推进的持久评论家相结合的PMD。对于有限折扣MDPs,我们建立了精确坐标式Bellman更新的值函数的全局线性收敛性,其中演员步长可为任意正常数,评论家初始化可为任意有限值。证明结合了基于预解式的辅助分布、衰减的Bellman违规校正以及由逆坐标权重加权的势函数。随后,我们在单一离策略马尔可夫轨迹下,研究了具有一般强凸镜像映射的随机TD-PMD。通过适当选择的常数步长和有限批次TD更新,该方法在$\widetilde{O}(1/((1-\gamma)^5 \widetilde{\sigma}_b \epsilon))$次转移后达到期望值差距$\epsilon$。随机分析依赖于正则化器的轨迹-wise Lipschitz连续性,该连续性由顶点Bregman散度的均匀界推导得出,并结合了用于有符号评论家误差传播的访问加权预解估计,该估计产生了对行为覆盖度$\widetilde{\sigma}_b$的逆线性依赖。与许多先前正则化策略优化的保证相比,我们的样本复杂度保证无需轨迹重置、生成模型访问或嵌套策略评估循环。数值结果与理论收敛分析一致。
英文摘要
Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize and arbitrary finite critic initialization. The proof combines a resolvent-based auxiliary distribution with a decaying Bellman-violation correction and a potential weighted by inverse coordinate weights. We then study stochastic TD-PMD with general strongly convex mirror maps under a single off-policy Markov trajectory. With suitably chosen constant stepsizes and a finite-batch TD update, the method achieves an expected value gap of $ε$ after $\widetilde{O}(1/((1-γ)^5 \widetildeσ_b ε))$ transitions. The stochastic analysis relies on the trajectory-wise Lipschitz continuity of the regularizer, derived from uniform bounds on vertex Bregman divergences, together with a visitation-weighted resolvent estimate for signed critic-error propagation that yields an inverse-linear dependence on behavior coverage $\widetildeσ_b$. In contrast to many prior guarantees for regularized policy optimization, our sample-complexity guarantee holds without trajectory resets, generative-model access, or nested policy-evaluation loops. Numerical results are consistent with the theoretical convergence analysis.