策略优化的良性非凸景观:一般状态与动作空间上的无限时域折扣马尔可夫决策过程
Benign Nonconvex Landscape for Policy Optimization: Infinite-Horizon Discounted MDPs with General State and Action Spaces
浏览论文内容
中文总结 AI 辅助
本文针对一般状态与动作空间的无限时域折扣MDP,提出弱条件排除次优平稳点并建立PLK条件,验证于两类运营模型,给出策略梯度方法首次非渐近收敛率。
中文摘要 AI 辅助
我们研究了在结构化平稳策略类下,具有一般状态与动作空间的无限时域折扣马尔可夫决策过程(MDPs)的优化景观。一种用于建立策略梯度方法全局收敛保证的一般加权策略迭代方法,要求在每一个策略下对加权策略改进封闭,而这一性质即使在策略类包含最优策略时也可能不成立。为解决此问题,我们提出了较弱的条件,以保证不存在次优平稳点,并建立了具有有限集中性系数的策略梯度目标的Polyak--Lojasiewicz--Kurdyka(PLK)条件。我们还从在每个状态成立的策略改进界建立了PLK条件,无需集中性假设。当共同的常设假设对相同的策略类和参数域成立时,我们的一般结果涵盖了早期框架所覆盖的设置。我们进一步验证了所提条件在两个运营模型上的适用性:具有马尔可夫调制需求的库存系统和随机现金余额问题。对于这两个模型,贝尔曼方程在动作变量上产生Q值函数的近似凸性,偏差由一阶平稳性度量控制。这些估计建立了指数一阶,并在额外曲率假设下建立指数二阶PLK条件,这些条件连同策略梯度的Lipschitz连续性,分别意味着使用精确策略梯度的投影梯度下降具有$\mathcal{O}(1/\epsilon)$的迭代复杂度和线性收敛。据我们所知,我们首次提供了使用策略梯度方法求解具有马尔可夫调制需求的无限时域折扣库存系统和随机现金余额问题的非渐近收敛速率。
英文摘要
We study the optimization landscape for infinite-horizon discounted Markov decision processes (MDPs) with general state and action spaces under structured stationary policy classes. A general weighted policy-iteration approach to establishing global convergence guarantees for policy gradient methods requires closure under weighted policy improvement at every policy, a property that may fail even when the policy class contains an optimal policy. To address this issue, we propose weaker conditions that guarantee the absence of suboptimal stationary points and establish the Polyak--Lojasiewicz--Kurdyka (PLK) condition for the policy gradient objective with a finite concentrability coefficient. We also establish the PLK condition from a policy-improvement bound that holds at every state, without a concentrability assumption. Our general results encompass settings covered by the earlier framework when the common standing assumptions hold for the same policy class and parameter domain. We further verify our proposed conditions for two operations models: inventory systems with Markov-modulated demand and stochastic cash-balance problems. For both models, the Bellman equation yields approximate convexity of the Q-value functions in the action variable, with deviations controlled by the first-order stationarity measure. These estimates establish exponent-one and, under additional curvature assumptions, exponent-two PLK conditions, which, together with Lipschitz continuity of the policy gradient, imply an $\mathcal{O}(1/ε)$ iteration complexity and linear convergence, respectively, for projected gradient descent using exact policy gradients. To the best of our knowledge, we provide the first non-asymptotic convergence rates for solving infinite-horizon discounted inventory systems with Markov-modulated demand and stochastic cash-balance problems using policy gradient methods.