发表机构
Delft University of Technology; University of Oxford(代尔夫特理工大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对带无限时域联合机会约束的随机最优控制问题,通过状态增广转化为约束马尔可夫决策过程,利用对偶性提出对偶上升算法,并结合离线学习近似值函数以降低计算复杂度,经数值示例验证其有效性。
AI 中文摘要
本文研究带无限时域联合机会约束的随机最优控制问题。通过合适的状态增广,将原问题重新表述为带约束的马尔可夫决策过程,其中代价函数和约束函数均具有可加结构。随后证明该表述具有强对偶性,从而可将问题重新表述为拉格朗日对偶框架下的等价无约束问题。提出对偶上升算法求解所得问题,并证明其收敛到增广状态空间上定义的确定性马尔可夫策略,该策略兼具最优性与可行性。为适配连续状态-输入空间,提出专用学习算法在离线训练场景下近似值函数,显著降低在线控制阶段的计算复杂度。最后通过数值示例测试所提方法,与在线预测控制方法在性能和计算复杂度方面对比,验证其有效性。
英文摘要
In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. We then prove that this formulation enjoys strong duality, thereby enabling us to reformulate the problem as an equivalent unconstrained one in the Lagrange dual framework. We propose a dual-ascent algorithm to solve the resulting problem and show that it converges to a deterministic Markov policy defined over the augmented state space that is both optimal and feasible. To accommodate continuous state-input spaces, we propose a dedicated learning algorithm to approximate the value function in an offline training setting, thereby significantly reducing the computational complexity of the online control phase. We then test our approach on a numerical example and demonstrate its effectiveness compared to online predictive control methods in terms of performance and computational complexity.
CommentsSubmitted to IEEE Transactions on Automatic Control