退化扩散随机控制的神经反馈近似:误差估计与数值分析
Neural feedback approximation for stochastic control with degenerate diffusions: error estimates and numerical analysis
浏览论文内容
中文总结 AI 辅助
研究有限时间随机最优控制问题,通过神经网络反馈映射近似时间离散公式,证明误差收敛估计,涵盖退化扩散等,通过数值实验说明方法并区分主要误差源。
中文摘要 AI 辅助
我们研究有限时间随机最优控制问题,并通过神经网络反馈映射上的直接策略学习问题来近似由此产生的时间离散公式。我们证明了在平均意义下,时间离散值与近似优化的神经策略所诱导的值之间误差的定量收敛估计。该界限区分了近最优反馈策略的近似、随机轨迹在紧集上的定位以及训练中的优化容差。分析不需要转移密度假设,并在统一框架中涵盖可能的退化扩散和确定性受控动力学。针对退化随机径向目标问题(a degenerate stochastic radial target problem)、哈密顿 - 雅可比 - 贝尔曼基准(a Hamilton--Jacobi--Bellman benchmark)和储气问题(a gas storage problem)提供了数值实验,说明了该方法并区分了主要误差来源:时间离散化、对分段常数策略的限制、神经网络近似和蒙特卡罗评估。
英文摘要
We study finite-horizon stochastic optimal control problems and approximate the resulting time-discrete formulation by a direct policy-learning problem over neural-network feedback maps. We prove a quantitative convergence estimate, in an averaged sense, for the error between the time-discrete value and the value induced by an approximately optimized neural policy. The bound separates the approximation of near-optimal feedback policies, the localization of stochastic trajectories on compact sets, and the optimization tolerance in training. The analysis does not require transition-density assumptions and covers possibly degenerate diffusions and deterministic controlled dynamics in a unified framework. Numerical experiments are provided for a degenerate stochastic radial target problem, a Hamilton--Jacobi--Bellman benchmark, and a gas storage problem, illustrating the approach and separating the main error sources: time discretization, restriction to piecewise-constant policies, neural-network approximation, and Monte Carlo evaluation.