arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32933cs.LG

结构化约束马尔可夫决策过程中的在线强化学习最后迭代保证

Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs

Nam Phuong Tran, Trinh Ha Mai Huynh, Tuyen Pham Le, Van-Truong Nguyen, Quan Nguyen, Long Tran-Thanh

AI总结:

针对结构化约束马尔可夫决策过程,提出统计高效的框架,分离正则化原始-对偶动力学与演员近似及统计误差,实现无模型最后迭代收敛,并在线性CMDP上验证稳定性。

AI中文摘要:

在安全关键应用中,部署使用单一策略,其性能和约束满足应直接成立,而非仅针对训练策略的平均或混合。这推动了约束强化学习中最后迭代保证的研究。近期进展已在精确梯度或表格型在线设置中建立了此类保证,但针对结构化大规模状态问题的可扩展结果仍然开放。我们开发了一个通用的、统计高效的框架,用于结构化约束马尔可夫决策过程(CMDPs)中的最后迭代收敛。我们的分析将正则化原始-对偶动力学的收缩与在线探索下策略评估中的演员近似和统计误差分离。这使得在结构化函数逼近下实现无模型的在线和离线学习成为可能:乐观策略评估避免了显式转移模型的构建,而紧凑的参数化演员避免了维护过去策略的混合或历史。我们将该框架实例化用于线性CMDPs和一般函数逼近,获得了依赖于表示形式的复杂度,并相比先前的乐观正则化原始-对偶分析改进了目标精度依赖性。我们进一步在合成线性CMDP上验证了理论预测的稳定效应:正则化方法表现出稳定的最后迭代行为,而其未正则化对应方法显示出更大的振荡。

英文摘要:

In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has established such guarantees in exact-gradient or tabular online settings, yet scalable results for structured large-state problems remain open. We develop a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs). Our analysis separates contraction of the regularised primal-dual dynamics from actor approximation and statistical errors in policy evaluation under online exploration. This enables model-free on- and off-policy learning with structured function approximation: optimistic policy evaluation avoids explicit transition-model construction, while a compact parametric actor avoids maintaining mixtures or histories of past policies. We instantiate the framework for linear CMDPs and general function approximation, obtaining representation-dependent complexity and improved target-accuracy dependence over prior optimistic regularised primal-dual analyses. We further validate the stabilising effect predicted by our theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.

↑