AI 中文总结
研究有限时域马尔可夫决策过程中自然策略梯度,给出其在恒定步长和递增步长下的收敛保证,恒定步长下次线性收敛,递增步长下线性收敛,特定步长调度可达几何速率。
AI 中文摘要
自然策略梯度(NPG)是一种成熟的强化学习算法,是信任区域策略优化和近端策略优化等广泛使用方法的基础,二者在实证中都很成功。本文研究具有已知动态和与视界相关转移核的有限时域马尔可夫决策过程中的精确NPG。我们首次在此设置下为该算法提供有限时间收敛保证,考虑了恒定步长和递增步长情况。恒定步长时,经t次迭代后NPG以\(\mathcal{O}(H^{2}/t)\)的速率次线性收敛;递增步长时,算法以\(\mathcal{O}\left(\left(1-\frac{1}{\vartheta_\rho}\right)^t\right)\)的线性收敛速率收敛,特定步长调度能达到此几何速率。
英文摘要
Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success. In this paper, we study exact NPG in finite-horizon Markov Decision Processes with known dynamics and horizon-dependent transition kernels. We provide the first finite-time convergence guarantees for this algorithm in this setting, for which we consider both constant and increasing step size regimes. With a constant step size $η_t=η$, we prove that NPG converges sublinearly with a rate of $\mathcal{O}(H^{2}/t)$ after $t$ iterations, where $H$ is the horizon length. We also extend this constant step size analysis to linear MDPs in an exact population-projection oracle under a full support projection distribution, recovering the same sublinear rate as in the tabular setting. Furthermore, with increasing step sizes, we prove that this algorithm achieves a linear convergence rate of $\mathcal{O}\left(\left(1-\frac{1}{\vartheta_ρ}\right)^t\right)$ for a problem-dependent constant $\vartheta_ρ> 1$, and the horizon-only robust schedule of the form $η_t=η_0(H/(H-1))^t$ where $η_0>0$ and $H \geq 2$, attains this same geometric rate.
Comments33 pages, 3 figures (each with two subfigures), 1 table. Accepted at the Reinforcement Learning Conference (RLC 2026)