arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于强化学习的马尔可夫非凸ADMM:超越光滑块的Bellman-预解稳定性

Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks

Zhaojun Peng

arXiv 2609.36859首次发表:更新:

AI 中文总结

本研究揭示马尔可夫非凸ADMM中折扣Bellman预解算子可提供乘子稳定性,建立强化学习收敛性,并给出迭代与样本复杂度,证明覆盖的KKT点全局最优,提出可复用的结构原则。

AI 中文摘要

我们识别并研究了一种用于强化学习中马尔可夫非凸ADMM的结构性机制。以有限折扣MDP作为典型验证平台,我们证明折扣Bellman预解算子$(I-\gamma P_\pi)^{-1}$能够提供经典非凸ADMM分析中通常从指定光滑块获得的乘子稳定性。基于这一机制,我们首先在受控马尔可夫采样下建立了收敛性,随后在随机观测下利用经验Bellman替代函数(该函数联合表示随机残差及其雅可比矩阵)建立了收敛性。马尔可夫混合、初始化漂移、观测噪声和衰减偏差作为单一算子扰动进入分析,避免了无偏乘积和双重采样的要求。当扰动平方可和时,真实KKT残差几乎必然收敛到零。在有限条件四阶矩条件下,伴随迭代满足$\mathbb{E}[\widetilde G_{K+1}] \le A/T+(B/T)\sum_{k<T}m_k^{-1}$,对于总马尔可夫样本预算$N$,该式变为$O(T^{-1}+T/N)$,从而对平方KKT精度$\epsilon$给出$O(\epsilon^{-1})$迭代复杂度和$O(\epsilon^{-2})$样本复杂度。超越平稳性,折扣占用覆盖率给出直接表格策略的$J^\star-J(\pi)=O(\sqrt G)$,因此覆盖的精确KKT点是全局最优的,而逐状态的二次Bellman改进条件将该关系加强为$O(G)$。最后,非线性策略、投影Bellman和显式占用率公式展现出相同的算子可逆性、对偶表示和乘子稳定性链条。这支持将折扣算子可逆性作为原始-对偶强化学习中可复用的结构原则。

英文摘要

We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-γP_π)^{-1}$ can provide the multiplier stability that classical nonconvex ADMM analyses often obtain from a designated smooth block. Starting from this mechanism, we establish convergence under controlled Markov sampling and then under stochastic observations using an empirical Bellman surrogate that jointly represents the random residual and its Jacobian. Markov mixing, initialization drift, observation noise, and decaying bias enter as one operator perturbation, avoiding unbiased product and double sampling requirements. When the perturbations are square summable, the true KKT residual converges almost surely to zero. Under a finite conditional fourth moment condition, a companion iterate satisfies $ \mathbb{E}[\widetilde G_{K+1}] \le A/T+(B/T)\sum_{k<T}m_k^{-1}, $ which becomes $O(T^{-1}+T/N)$ for total Markov sample budget $N$, giving $O(ε^{-1})$ iteration complexity and $O(ε^{-2})$ sample complexity for squared KKT accuracy $ε$. Beyond stationarity, discounted occupancy coverage yields $J^\star-J(π)=O(\sqrt G)$ for direct tabular policies, so covered exact KKT points are globally optimal, while a statewise quadratic Bellman-improvement condition sharpens the relation to $O(G)$. Finally, nonlinear policy, projected Bellman, and explicit occupancy formulations exhibit the same chain of operator invertibility, dual representation, and multiplier stability. This supports discounted operator invertibility as a reusable structural principle for primal-dual reinforcement learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑