arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38395math.OCcs.LGcs.ROcs.SYeess.SYmath.PR

粗糙路径与Itô框架下随机最优控制的统一最优性条件

Unified Optimality Conditions for Stochastic Optimal Control in the Rough Path and Itô Frameworks

Thomas Lew

首次发表
浏览论文内容

中文总结 AI 辅助

本文建立了Itô与粗糙路径框架下随机最优控制PMP条件的统一形式,通过条件期望桥接伴随方程,并应用于生成模型微调与反馈控制。

中文摘要 AI 辅助

随机微分方程(SDEs)可以通过Itô微积分和粗糙路径理论进行研究。对于随机最优控制,这两种框架给出了不同的庞特里亚金最大值原理(PMP)最优性条件,分别涉及前向-后向随机微分方程(FBSDEs)或粗糙微分方程。我们证明了Itô和粗糙PMP的伴随方程通过条件期望$p_t^{\text{Itô}}=\mathbb{E}[p_t^{\text{rough}} \mid \mathcal{F}_t]$相联系,其中$\mathcal{F}_t$表示在时间$t$可用的信息。首先,我们针对具有适应控制的问题推导了一个不使用FBSDEs的粗糙随机PMP。其证明通过考虑随机针状变分,将确定性控制上的粗糙随机PMP进行了扩展。其次,我们利用Itô-Stratonovich转换公式以及前向切向SDE与后向伴随SDE之间的对偶恒等式,推导了一个连接Itô和粗糙PMP的统一PMP。作为第一个应用,我们重新推导了用于微调生成模型的伴随匹配方法。作为第二个应用,我们针对一类反馈问题提出了一种间接打靶法。总体而言,这些结果为连接随机最优控制的两种流行框架提供了一座新的条件桥梁。

英文摘要

Stochastic differential equations (SDEs) can be studied via Itô calculus and rough path theory. For stochastic optimal control, these two frameworks give distinct Pontryagin Maximum Principle (PMP) optimality conditions with forward-backward SDEs (FBSDEs) or rough differential equations. We show that the adjoint equations of the Itô and rough PMPs are connected via the conditional expectation $p_t^{\text{Itô}}=\mathbb{E}[p_t^{\text{rough}} \mid \mathcal{F}_t]$, where $\mathcal{F}_t$ represents information available at time $t$. First, we derive a rough stochastic PMP for problems with adapted controls that does not use FBSDEs. Its proof extends the rough stochastic PMP over deterministic controls by considering stochastic needle variations. Second, we derive a unified PMP connecting the Itô and rough PMPs, using Itô-Stratonovich conversion formulas and duality identities between the forward tangent and backward adjoint SDEs. As a first application, we rederive the adjoint matching method for fine-tuning generative models. As a second application, we propose an indirect shooting method for a class of feedback problems. Overall, these results give a new conditional bridge connecting two popular frameworks for stochastic optimal control.

发表机构

  • Toyota Research Institute(丰田研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑