AI 中文总结
研究有限折扣马尔可夫决策过程中任意恒定步长的非正则化策略镜像下降的策略收敛性,基于可分解镜像映射证明其收敛到最优策略,涵盖多种镜像映射,还分析了收敛行为及建立了局部收敛速率。
AI 中文摘要
我们研究了有限折扣马尔可夫决策过程中具有任意恒定步长的非正则化策略镜像下降(PMD)的策略收敛性。我们关注形式为\(h(p)=\sum_a \psi(p(a))\)的可分解镜像映射,其中\(\psi\)满足标准的勒让德型假设。在这些条件下,我们证明了即使最优策略集不是单元素集,由PMD生成的策略序列在策略域中也会收敛到一个极限最优策略。这个结果涵盖了一大类常用的镜像映射,包括投影Q上升背后的平方欧几里得镜像映射、softmax自然策略梯度背后的负香农熵、Tsallis熵、赫林格映射和费米 - 狄拉克熵。尽管之前已经针对特定的PMD实例或正则化变体(如同伦PMD)建立了策略收敛性,但据我们所知,这是第一个在一般可分解镜像映射和任意恒定步长下关于非正则化PMD的系统且统一的策略收敛理论。我们的分析进一步表明,收敛行为由\(\psi\)在\(0\)和\(1\)处的可微性决定,导致不同的行为,包括有限时间终止、渐近收敛和依赖于MDP的二分法。当\(\psi\)具有严格正的有限曲率的二次连续可微时,我们进一步为渐近收敛情况建立了局部策略收敛速率,涵盖了上述标准镜像映射。
英文摘要
We study the policy convergence of unregularized policy mirror descent (PMD) with arbitrary constant step sizes for finite discounted Markov decision processes. We focus on decomposable mirror maps of the form $h(p)=\sum_a ψ(p(a))$, where $ψ$ satisfies standard Legendre-type assumptions. Under these conditions, we prove that the policy sequence generated by PMD converges in the policy domain to a limiting optimal policy, even when the optimal policy set is not a singleton. This result covers a broad class of commonly used mirror maps, including the squared Euclidean mirror map underlying projected Q-ascent, the negative Shannon entropy underlying softmax natural policy gradient, Tsallis entropy, the Hellinger mapping, and the Fermi-Dirac entropy. Although policy convergence has been established previously for specific PMD instances or for regularized variants such as homotopic PMD, to the best of our knowledge, this is the first systematic and unified policy convergence theory for unregularized PMD under general decomposable mirror maps and arbitrary constant step sizes. Our analysis further reveals that the convergence behavior is governed by the differentiability of $ψ$ at $0$ and $1$, leading to different behaviors, including finite-time termination, asymptotic convergence, and an MDP-dependent dichotomy. When $ψ$ is twice continuously differentiable with strictly positive finite curvature, we further establish local policy convergence rates for the asymptotic convergence cases, covering the standard mirror maps mentioned above.