arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19587cs.LG

单循环熵正则化自然演员-评论家算法的无正则化收敛性

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

Zhiqiang Tan

首次发表
浏览论文内容

中文总结 AI 辅助

本文分析单循环熵正则化自然演员-评论家算法,通过指数平移机制实现无正则化收敛加速,在随机和确定性场景中均突破传统统计屏障。

中文摘要 AI 辅助

尽管熵正则化被广泛用于稳定和加速自然策略梯度方法,但它对无正则化目标函数实现更快收敛速率的能力仍未得到充分探索。现有分析通常依赖双循环架构并引入线性熵惩罚项。为弥合理论与实践之间的差距,我们在兼容线性函数近似下分析了一种单循环熵正则化自然演员-评论家算法。通过训练非中心化评论家,即使训练策略趋近于确定性且费舍尔信息矩阵退化,我们的评论家跟踪仍能保持稳定。我们关注优化景观的两种主要场景:随机场景中,我们将耦合的演员-评论家更新融合为联合李雅普诺夫递推;确定性场景中,我们转向策略镜像下降框架以规避欧几里得几何的崩溃。通过利用无正则化马尔可夫决策过程中的正最小动作间隙,我们引入指数平移机制,将正则化间隙映射到无正则化间隙,直至指数衰减尾项。通过调整固定温度,我们的算法实现了加速的无正则化收敛速率,近似误差项除外:随机场景中为$\tilde{\tau}(T_{total}^{-1})$,确定性场景中平均迭代为$\tilde{\tau}(T_{total}^{-2/3})$、最后一次迭代为$\tilde{\tau}(T_{total}^{-1/3})$。此处,$T_{total}$表示随机评论家更新(或蒙特卡洛回合)的总次数。此外,在表格型设置中,我们的正动作间隙分析得到平均迭代速率$\tilde{\tau}(T_{total}^{-2/3})$,超过了无正动作间隙时适用的最坏情况统计屏障$\tau(T_{total}^{-1/2})$。

英文摘要

While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty. To bridge the gap between theory and practice, we analyze a single-loop, entropy-regularized Natural Actor-Critic algorithm under compatible linear function approximation. By training an uncentered critic, our critic tracking can remain stable even as the training policy approaches determinism and the Fisher information matrix degenerates. We focus on two primary regimes for the optimization landscape: a Stochastic Regime, where we fuse coupled actor-critic updates into a joint Lyapunov recurrence, and a Deterministic Regime, where we pivot to a Policy Mirror Descent framework to circumvent the collapse of Euclidean geometry. By exploiting a positive Minimal Action Gap in the unregularized Markov decision process, we introduce an Exponential Translation mechanism that maps the regularized gap to the unregularized one up to an exponentially decaying tail. By tuning the fixed temperature, our algorithm achieves accelerated unregularized convergence rates, up to approximation-error terms: $\tilde{\mathcal{O}}(T_{total}^{-1})$ in the Stochastic Regime, and $\tilde{\mathcal{O}}(T_{total}^{-2/3})$ for the average iterate alongside $\tilde{\mathcal{O}}(T_{total}^{-1/3})$ for the last iterate in the Deterministic Regime. Here, $T_{total}$ denotes the total number of stochastic critic updates (or Monte Carlo rollouts). Furthermore, in the tabular setting, our positive-action-gap analysis yields a $\tilde{\mathcal{O}}(T_{total}^{-2/3})$ average-iterate rate, surpassing the $\mathcal{O}(T_{total}^{-1/2})$ worst-case statistical barrier that applies without a positive action margin.

发表机构

  • Rutgers University(罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑