发表机构
UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对非马尔可夫环境下霍克斯驱动随机微分方程的随机控制问题,提出Hawkes-CT DDPG算法,通过马尔可夫化近似实现无模型学习,并在三类核上与离散时间强化学习方法对比。
AI 中文摘要
我们在非马尔可夫环境下,采用机器学习算法研究多变量霍克斯驱动随机微分方程的随机控制问题。由于霍克斯强度的记忆具有路径依赖性,该问题在马尔可夫核之外无法归入经典随机控制理论范畴。我们首先开发有限维马尔可夫化程序与算法,用指数核混合模型近似多变量霍克斯过程,证明霍克斯过程、其强度及该问题价值的马尔可夫化近似,会收敛至原始非马尔可夫过程及原问题的价值。接着,我们在该马尔可夫化近似问题上构建连续时间确定性策略梯度学习方法,命名为Hawkes-CT DDPG。我们提出一种无模型算法,用于解决非马尔可夫霍克斯驱动优化问题,该算法仅需观测过程的事件时间、随机微分方程解的实现值及选定的一组衰减滤波器,而霍克斯核系数保持未知。我们在三种不同类型的核(简单指数核、埃尔朗核、幂律核)下,将我们的连续时间强化学习方法Hawkes-CT DDPG与离散时间强化学习技术进行对比。
英文摘要
We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. We first develop a finite-dimensional Markovianization procedure and algorithm to approximate multivariate Hawkes processes with mixtures of exponential kernels. We prove the convergence of the Markovianized approximation of the Hawkes process, its intensity, and the value of the problem to the original non-Markovian processes and the value of the primal problem. We then formulate continuous-time deterministic policy gradient learning on the Markovianized approximation of the problem, called Hawkes-CT DDPG. We propose a model-free algorithm to solve the non-Markovian Hawkes-driven optimization by observing only the event times of the process, the realization of the solution to the SDE, and a chosen set of decay filters, while the Hawkes kernel coefficients remain unknown. We compare our continuous time reinforcement learning Hawkes-CT DDPG method with discrete time reinforcement learning techniques under three different types of kernels: simple exponential, Erlang, and power-law kernels.