arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

机会约束协方差转向的马尔可夫策略最优性

On the Optimality of Markovian Policies for Chance-Constrained Covariance Steering

Naoya Kumagai, Kenshiro Oguri

arXiv 2608.22589首次发表:更新:

AI 中文总结

本文证明机会约束协方差转向问题中,依赖历史的最优策略可无损转化为马尔可夫策略,拓展分析至输出反馈与风险价值代价代理问题。

AI 中文摘要

许多针对有限时段随机最优控制(包括协方差转向)的研究,将控制策略参数化为状态历史仿射形式。这种参数化可实现凸重构,从而得到可处理的求解方法。但依赖前序状态的必要性尚未得到充分证明:这种依赖是必要的,还是仅为凸重构的人为产物?本文证明其为可无损移除的人为产物。给定状态历史仿射形式的最优解,我们构造一个对当前状态仿射的确定性马尔可夫策略。结果表明,即使针对带有广泛常用状态与控制安全约束的协方差转向问题,所合成的马尔可夫策略也几乎必然产生与依赖历史策略相同的控制动作,进而得到相同的状态轨迹、代价与矩。因此,依赖历史形式的每个最优解都可进行无损马尔可夫变换。从几何角度看,依赖历史形式为凸性提升了策略空间,其最优解可投影回马尔可夫策略空间。我们将分析扩展至输出反馈及风险价值代价的凸上界代理问题。

英文摘要

Many studies on finite-horizon stochastic optimal control, including covariance steering, parameterize control policies as state-history-affine. This parameterization enables a convex reformulation, thereby yielding a tractable solution method. However, the necessity of dependence on previous states has not been well established. \textit{Is this dependence necessary, or merely an artifact of the convex reformulation?} We show that it is an artifact that can be removed losslessly. Given an optimal solution of the state-history-affine formulation, we construct a deterministic Markovian policy which is affine in the current state. We show that, even for the covariance steering problem with a broad class of commonly used state and control safety constraints, the synthesized Markovian policy almost surely produces the same control actions as the history-dependent policy and therefore the same state trajectories, cost, and moments. Thus, every optimum of the history-dependent formulation admits a lossless Markovian transformation. Geometrically, the history-dependent formulation lifts the policy space for convexity, and its optimal solution can be projected back to the Markovian policy space. We extend the analysis to output feedback and a convex upper-bounding surrogate for value-at-risk costs.

Comments12 pages, 2 figures. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑