arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38745math.OC

连续时间与状态的无模型强化学习:一种随机最大值原理方法

Model-free Reinforcement Learning for Continuous Time and State: A Stochastic Maximum Principle Approach

  • Xidian University(西安电子科技大学)
  • The Hong Kong Polytechnic University(香港理工大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Lijun Bo, Yijie Huang, Jingfei Wang

AI总结:

本文提出一种基于随机最大值原理的无模型强化学习算法,用于连续时间随机控制问题,通过直接学习哈密顿梯度实现策略梯度优化,并证明其收敛性及与现有方法的等价性。

AI中文摘要:

本文针对具有连续状态和动作空间的连续时间随机控制问题,基于随机最大值原理开发了一种无模型强化学习(RL)算法。对于参数化的马尔可夫策略,我们建立了伴随倒向随机微分方程解耦场的存在性,这使得哈密顿梯度能够表示为时间和状态的确定性函数。然后,我们直接从数据中学习该哈密顿梯度,并将其纳入具有不精确梯度的策略梯度方案中。我们建立了所得策略梯度算法的收敛性,并给出了相应的误差估计。在学习率、探索参数和逼近误差满足适当条件的情况下,目标值收敛,且策略梯度的$L^2$范数渐近消失。我们进一步证明了所提出的基于SMP的策略梯度表示与基于优势率函数的现有连续时间确定性策略梯度表示是等价的。数值实验验证了所提出的RL算法的有效性和效率。

英文摘要:

This paper develops a model-free reinforcement learning (RL) algorithm based on the stochastic maximum principle for continuous-time stochastic control problems with continuous state and action spaces. For a parameterized Markovian policy, we establish the existence of the decoupling field for the adjoint backward stochastic differential equation, which allows the Hamiltonian gradient to be represented as a deterministic function of time and state. We then learn this Hamiltonian gradient directly from data and incorporate it into a policy gradient scheme with inexact gradients. We establish the convergence of the resulting policy gradient algorithm and establish the corresponding error estimates. Under suitable conditions on learning rates, exploration parameters, and approximation errors, the objective values converge and the $L^2$-norm of the policy gradient vanishes asymptotically. We further show that the proposed SMP-based policy gradient representation is equivalent to the existing continuous-time deterministic policy gradient representation based on the advantage-rate function. The effectiveness and efficiency of the proposed RL algorithm are demonstrated by numerical experiments.

↑