双时间尺度Actor-Critic算法的浓度界
A Concentration Bound for Two-Timescale Actor-Critic Algorithm
- IISc Bangalore(印度科学研究所班加罗尔)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对长期平均奖励设置下的双时间尺度actor-critic算法,推导了均匀全时浓度界,证明actor参数以高概率进入并保持在安全区域,误差随更新次数减小。
AI中文摘要:
近年来,大量研究工作致力于为双时间尺度actor-critic算法建立渐近和非渐近收敛保证,其中actor递归在比critic递归更慢的时间尺度上运行。本文在长期平均奖励设置下,推导了带函数逼近的actor-critic算法的均匀全时浓度界。该界有助于我们以高概率分析actor参数的行为。我们证明,在有限时间之后,actor参数以高概率进入安全区域并此后保持在该区域内。具体而言,以至少$1-\epsilon_1-\epsilon_2$的概率,对于所有$k\geq n_0$且$n_0$足够大,actor误差$\Vert \theta_k-\theta^{*}\Vert$为$O\left(\frac{n_0^{3/4}}{k}\frac{1}{\sqrt{\epsilon_2}}+\left(\frac{1}{n_0}\right)^{1/4}\log^{1/4}\left(\frac{1}{\epsilon_1}\right)+\left(\frac{1}{n_0}\right)^{1/4}\right)$。我们还展示了实验结果,表明上述actor误差随actor参数更新次数的增加而减小。
英文摘要:
Significant research effort has been directed in recent years towards establishing both asymptotic and non-asymptotic convergence guarantees for two-timescale actor--critic algorithms, where the actor recursion is run on a slower timescale than the critic recursion. This work derives a uniform all-time concentration bound for the actor--critic algorithm with function approximation in the long-run average-reward setting. This bound helps us analyze the behavior of the actor parameter with high probability. We show that, after some finite time, the actor parameter enters a safe region and remains within it thereafter with high probability. Specifically, with probability at least $1-ε_1-ε_2$, the actor error $\Vert θ_k-θ^{*}\Vert$ is $O\left(\frac{n_0^{3/4}}{k}\frac{1}{\sqrt{ε_2}}+\left(\frac{1}{n_0}\right)^{1/4}\log^{1/4}\left(\frac{1}{ε_1}\right)+\left(\frac{1}{n_0}\right)^{1/4}\right)$ for all $k\geq n_0$ and sufficiently large $n_0$. We also present experimental results demonstrating that the aforementioned actor error diminishes with the number of actor-parameter updates.