arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

平均回报约束马尔可夫决策过程(CMDP)中实现阶最优神经演员-评论家的分层多级蒙特卡洛方法

Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs

Ankur Naskar, Vaneet Aggarwal

arXiv 2607.28390首次发表:更新:

AI 中文总结

该研究针对平均回报CMDP中神经评论家的阶最优收敛问题,提出分层MLMC神经评论家,开发出实现阶最优的原始-对偶自然演员-评论家算法,无需知晓混合时间。

AI 中文摘要

约束马尔可夫决策过程(CMDP)为安全关键应用中的强化学习提供了自然框架,智能体在最大化长期回报的同时需满足长期约束。尽管带线性评论家的原始-对偶演员-评论家方法已得到充分研究,但将阶最优收敛保证扩展到平均回报CMDP中的神经评论家仍是待解决问题。核心挑战在于神经评论家估计中存在根本性的偏差-代价权衡:在神经正切核(NTK)分析下,大幅降低评论家偏差会显著增加评论家优化代价,阻碍原始-对偶框架实现阶最优收敛。我们通过引入分层多级蒙特卡洛(MLMC)神经评论家解决该瓶颈,该方法在轨迹采样和评论家优化中同时执行去偏。所得估计器仅需对数级期望样本代价即可达到长评论家优化运行的偏差。基于此估计器,我们开发了原始-对偶自然演员-评论家算法,其最优性间隙和约束违反均达到$\tilde{O}(T^{-1/2})$阶。这为具有一般策略参数化和神经评论家的无限时域平均回报CMDP建立了首个阶最优收敛保证,且无需知晓底层混合时间,该结果在无约束场景下同样具有创新性。

英文摘要

Constrained Markov Decision Processes (CMDPs) provide a natural framework for reinforcement learning in safety-critical applications, where agents maximize long-term reward while satisfying long-term constraints. Although primal-dual actor-critic methods with linear critics are well understood, extending order-optimal convergence guarantees to neural critics in average-reward CMDPs has remained open. The main challenge is a fundamental bias-cost trade-off in neural critic estimation: under Neural Tangent Kernel (NTK) analysis, reducing critic bias substantially increases critic optimization cost, preventing order-optimal convergence in the primal-dual framework. We resolve this bottleneck by introducing a hierarchical Multilevel Monte Carlo (MLMC) neural critic that performs debiasing simultaneously across trajectory sampling and critic optimization. The resulting estimator attains the bias of a long critic optimization run with only logarithmic expected sample cost. Building on this estimator, we develop a primal-dual Natural Actor-Critic algorithm that achieves both an optimality gap and a constraint violation of order $\tilde{O}(T^{-1/2})$. This establishes the first order-optimal convergence guarantees for infinite-horizon average-reward CMDPs with general policy parameterization and neural critics, while eliminating the need to know the underlying mixing time. Our results are novel even in the unconstrained setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑