发表机构
Florida State University; Worcester Polytechnic Institute(佛罗里达州立大学; 伍斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对带非齐次同步耦合的随机控制问题,提出连续策略-价值迭代框架,避免每次迭代求解偏微分方程,证明其指数收缩性,经数值实验验证其策略改进性与收敛性。
AI 中文摘要
针对无限时间跨度随机控制问题,我们提出了一种连续策略-价值迭代框架:策略通过非齐次哈密顿量梯度驱动的朗之万型动力学演化,同时价值函数随最优策略迭代更新,避免每次迭代直接求解对应的偏微分方程。在合适的正则性与容许性假设下,该迭代方案具备策略改进性;我们通过非齐次同步耦合,在期望利普希茨条件下证明了联合策略-价值动力学在Wasserstein-2距离下的指数收缩性,并验证了线性二次问题满足这些条件。对线性二次模型及不完全市场因子模型中的最优消费-投资问题的数值实验,展示了策略改进性、向最优策略的收敛性,以及衰减探索噪声的效果。
英文摘要
We propose a continuous policy-value iteration framework for infinite horizon stochastic control problems, in which the policy evolves through Langevin type dynamics driven by the gradient of the inhomogeneous Hamiltonian, while the value function is updated simultaneously with optimal policy iterations, avoiding directly solving its corresponding PDE at each iteration. The iteration scheme enjoys policy improvement under suitable regularity and admissibility assumptions. We show the exponential contraction of the joint policy value dynamics in Wasserstein-2 distance under the Lipschitz condition in expectation via the inhomogeneous synchronous coupling and verify these conditions for linear quadratic problems. Numerical experiments on linear quadratic models and the optimal consumption-investment problems in an incomplete market factor model illustrate policy improvement, convergence toward optimal policies, and the effect of decaying exploration noise.