(廉价的)随机策略梯度在线性二次调节器中以高概率收敛
(Cheap) Stochastic Policy Gradient Converges with High Probability for Linear Quadratic Regulator
浏览论文内容
中文总结 AI 辅助
研究随机策略梯度方法在线性二次调节器问题中的高概率收敛性,提出廉价方法,每次迭代仅需少量交互,且交互次数对置信度呈多对数依赖,具有推广潜力。
中文摘要 AI 辅助
我们研究了将普通随机策略梯度方法应用于线性二次调节器(LQR)问题的收敛性。该方法在以下意义上是廉价的:(1)每次迭代仅需与环境进行$\tilde{O}(1)$次交互,因此允许频繁的策略改进步骤;(2)为确保整个过程中的稳定性并以概率$1-\delta$收敛到$\epsilon$-最优策略,仅需$O(\mathtt{Polylog}(1/\delta)/\epsilon)$次交互。据我们所知,这似乎是首次有随机无模型策略优化方法在LQR中以高概率收敛,且每次迭代计算量为$\tilde{O}(1)$,对置信水平的依赖为多对数级别。本文提出的收敛分析不依赖于LQR的具体特性,因此可能推广到更广泛的问题类别。
英文摘要
We study the convergence of the vanilla stochastic policy gradient method applied to the linear quadratic regulator (LQR) problem. The method is cheap in the following sense: (1) at each iteration only $\tilde{O}(1)$ interactions with the environment are needed, therefore allowing frequent policy improvement steps, and (2) to ensure stability throughout and convergence to an $ε$-optimal policy with probability $1-δ$, only $O(\mathtt{Polylog}(1/δ)/ε)$ interactions are needed. To the best of our knowledge, this appears to be the first time that a stochastic model-free policy optimization method for LQR converges with high probability with $\tilde{O}(1)$ per-iteration computation and polylogarithmic dependence on the confidence level. The convergence analysis presented here is agnostic to LQR specifics and hence could be potentially generalized to a broader class of problems.
发表机构
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。