AI 中文总结
本文针对带模式切换的无限时段随机线性二次最优控制问题,提出仅从在线状态轨迹学习最优控制器的无模型Q学习算法,通过理论证明与数值案例验证了策略的等价性、系统稳定性及算法收敛性。
AI 中文摘要
本文研究无限时段连续时间带模式切换的随机线性二次最优控制问题。我们提出从基于模型的设计转向采用自适应动态规划方法,专门开发仅从在线状态轨迹数据学习最优控制器的在线策略和离线策略Q学习算法。本文的理论核心包括完整证明在线与离线策略架构的等价性,以及建立闭环系统稳定性和算法收敛到最优解的严格分析。为便于计算,我们使用向量化和克罗内克积代数实现这些算法。数值案例研究验证了理论结果,清晰展示了所提出的无模型控制策略的运行有效性和实际可行性。
英文摘要
This paper addresses infinite-horizon continuous-time stochastic linear quadratic optimal control problems with regime switching. We propose a paradigm shift from model-based design by adopting an adaptive dynamic programming approach, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data. The theoretical core of our work consists of a complete proof of the equivalence between the on- and off-policy architectures, alongside a rigorous analysis establishing the stability of the closed-loop system and the convergence of the algorithms to the optimal solution. For computational tractability, we implement these algorithms using vectorization and Kronecker product algebra. The theoretical results are corroborated by numerical case studies that clearly demonstrate the operational effectiveness and practical feasibility of the proposed model-free control strategy.
Comments21 pages, 1 figure