arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35466cs.LGmath.OCstat.ML

镜像下降软演员-评论家算法分析

An analysis of Mirror-Descent Soft Actor-Critic

Denis Zorba, Michal Valko

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明镜像下降软演员-评论家算法在连续动作空间中的收敛性,推导演员目标强凸光滑条件,获得O(N^(-1/5))最优迭代收敛率,并指出镜像下降步长控制目标漂移优于吉布斯目标。

中文摘要 AI 辅助

软演员-评论家(SAC)算法被广泛用于连续动作空间的熵正则化强化学习,实际实现中仅对不断变化的目标执行少量演员步骤。本文证明了当目标策略由策略镜像下降产生时的收敛性保证,并将其与经典吉布斯目标进行比较。我们推导了演员目标函数强凸性和光滑性的充分条件,这些条件通过勒让德微分算子由Q函数估计的曲率刻画,并建立了在演员和评论家近似误差下达到O(N^(-1/5))的最优迭代有限时间收敛率。此外,镜像下降步长λ直接控制目标漂移,从而影响演员跟踪误差,而类似的吉布斯界包含一个非消失的跟踪项。

英文摘要

Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare it with the classical Gibbs target. We derive sufficient conditions for the strong convexity and smoothness of the actor objective, characterised by the curvature of the $Q$-function estimate through the Legendre differential operator, and establish an $\mathcal{O}\!\left(N^{-\frac{1}{5}}\right)$ best-iterate finite-time convergence rate up to actor and critic approximation errors. Moreover, the mirror-descent step size $λ$ directly controls the target drift and hence actor tracking error, whereas the analogous Gibbs bound contains a non-vanishing tracking term.

发表机构

  • School of Mathematics, University of Edinburgh(爱丁堡大学数学学院)
  • INRIA(法国国家信息与自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑