arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习求解具有未知漂移和运行奖励的随机控制问题:理论、算法与收敛性

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence

Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou

arXiv 2609.14972首次发表:更新:

发表机构

University of Southern California; Columbia University(南加州大学; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对漂移和奖励未知的随机控制问题,提出基于探索性强化学习框架的策略迭代算法,利用概率表示求解最优价值函数与策略,并证明收敛性及数值验证。

AI 中文摘要

我们研究连续时间且可能高维的随机控制问题,其中漂移系数和运行奖励函数是未知的。由于这些模型基本要素缺失,我们采用Wang、Zariphopoulou和Zhou(2020)提出的探索性强化学习(RL)框架,该框架使用松弛控制和熵正则化。目标是开发具有理论依据、高效且可扩展的强化学习算法,以同时学习最优价值函数(其也求解探索性HJB方程)和最优探索性反馈控制策略。当扩散系数不包含控制时,我们基于仅依赖于原始动力学扩散部分的辅助状态过程,利用最优价值函数及其梯度的概率表示。通过对适当定义的映射及其不动点进行精细分析,我们引入了策略迭代算法并证明了其收敛性。我们通过各种数值示例展示了算法的性能。最后,我们研究了一个特殊的控制依赖扩散情形,其中需要用到Hessian矩阵的概率表示。

英文摘要

We study continuous-time and possibly high-dimensional stochastic control problems where drift coefficients and running reward functions are unknown. Due to these missing model primitives, we take the exploratory, reinforcement learning (RL) framework of Wang, Zariphopoulou, and Zhou(2020) with relaxed controls and entropy regularization. The objective is to develop theoretically grounded, efficient and scalable RL algorithms to learn both the optimal value functions (which also solve the exploratory HJB equation) and optimal exploratory feedback control policies. When the diffusion coefficients do not contain control, we employ probabilistic representations of both the optimal value function and its gradient based on an auxiliary state process depending only on the diffusion part of the original dynamics. With a delicate analysis on some properly defined mappings and their fixed points, this leads to the introduction of our policy iteration algorithms and their convergence. We demonstrate the performance of our algorithms through various numerical examples. Finally, we study a special control-dependent diffusion case where probability representation of the Hessian is called for.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑