无限时域随机最优控制中经验动态规划的渐近分析
Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control
AI总结:
研究无限时域随机最优控制问题,推导基于样本近似的统计极限定理,包括在不同条件下的中心极限定理,比较不同方法的渐近结果,还给出非唯一最优策略模型的极限定理,并通过实例说明理论。
AI中文摘要:
我们推导了离散时间下无限时域折扣随机最优控制问题基于样本近似的统计极限定理。首先在总体最优策略的唯一性条件下得到基于样本值函数的泛函中心极限定理,极限律是由类似动态规划原理的线性不动点方程刻画的均值为零的高斯过程。将这些渐近结果与基于样本策略优化的结果比较,表明极限方差可能不同。还推导了非唯一最优策略模型的极限定理,极限律可能非高斯。库存控制和可再生资源收获的应用说明了该理论。
英文摘要:
We derive statistical limit theorems for sample-based approximations of infinite-horizon discounted stochastic optimal control problems in discrete time. Our first result is a functional central limit theorem for the sample-based value function under a uniqueness-type condition on population optimal policies. The limiting law is a mean-zero Gaussian process characterized by a linear fixed point equation that resembles a dynamic programming principle. We compare these asymptotics with those obtained from sample-based policy optimization and illustrate that their limiting variances can be different. We also derive a limit theorem for models with nonunique optimal policies, where the limiting law may be non-Gaussian. Applications to inventory control and renewable harvesting illustrate the theory.