arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

方差缩减策略梯度的样本复杂度:更弱的假设与下界

Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds

Gabor Paczolay, Matteo Papini, Alberto Maria Metelli, Istvan Harmati, Marcello Restelli

arXiv 2610.03165首次发表:更新:

发表机构

Politecnico di Milano; Budapest University of Technology and Economics; University of Milan(米兰理工大学; 布达佩斯科技经济大学; 米兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出防御性策略梯度算法,在无重要性权重方差假设下达到$O(\epsilon^{-3})$样本复杂度,并通过广义黑盒模型下界证明其速率最优,优于普通策略梯度。

AI 中文摘要

基于重要性采样的若干REINFORCE方差缩减版本,在重要性权重方差的不切实际假设下,实现了改进的$O(\epsilon^{-3})$样本复杂度以找到$\epsilon$-稳定点。本文提出基于防御性重要性采样的\algo(防御性策略梯度)算法,在不对普通重要性权重方差作任何假设的情况下达到相同速率。我们还在一个隐藏状态和动作、允许参数相关奖励的广义黑盒策略优化模型中建立了下界。在该模型中,具有有界方差单策略反馈的最优速率为$\Theta(\epsilon^{-4})$,具有均方光滑耦合双策略反馈的最优速率为$\Theta(\epsilon^{-3})$。在标准策略正则性条件下,REINFORCE和\algo分别实现相应的预言机条件,并分别达到$O(\epsilon^{-4})$和$O(\epsilon^{-3})$的上界。尽管下界不直接适用于这些算法运行所在的经典MDP交互模型,但这种对应关系提供了预言机层面的证据,表明\algo的更快速率是最优的,并且与普通策略梯度真正分离。

英文摘要

Several variance-reduced versions of REINFORCE based on importance sampling achieve an improved $O(ε^{-3})$ sample complexity to find an $ε$-stationary point, under an unrealistic assumption on the variance of the importance weights. In this paper, we propose the \algo (Defensive Policy Gradient) algorithm, based on defensive importance sampling, which achieves the same rate without any assumption on the variance of ordinary importance weights. We also establish lower bounds in a generalized black-box policy-optimization model that hides states and actions and permits parameter-dependent rewards. In this model, the optimal rates are $Θ(ε^{-4})$ with bounded-variance one-policy feedback and $Θ(ε^{-3})$ with mean-square-smooth coupled two-policy feedback. Under standard policy-regularity conditions, REINFORCE and \algo realize the corresponding oracle conditions and attain the $O(ε^{-4})$ and $O(ε^{-3})$ upper bounds, respectively. Although the lower bounds do not apply directly to the classical MDP interaction model in which these algorithms operate, this correspondence provides oracle-level evidence that the faster rate of \algo is optimal and genuinely separated from that of vanilla policy gradient.

Journal refPaczolay, G., Papini, M., Metelli, A.M. et al. Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds. Mach Learn 113, 6475-6510 (2024)

DOI:10.1007/s10994-024-06573-4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑