arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散律的凸几何与随机控制的梯度流

Convex geometry of diffusion laws and gradient flows for stochastic control

Yupeng Bai, Louis-Pierre Chaintron, Zhenjie Ren, Songbo Wang

arXiv 2610.03487首次发表:更新:

发表机构

Université Evry Paris-Saclay; École Polytechnique Fédérale de Lausanne; Université Côte d’Azur(巴黎萨克雷大学埃夫里分校; 洛桑联邦理工学院; 蔚蓝海岸大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究弱随机控制的梯度型动力学,通过路径律的凸几何构造镜像流与自然梯度流,证明其收敛性、稳定性及有限粒子近似的指数收敛和O(1/N)误差。

AI 中文摘要

我们通过路径空间上受控扩散的律来研究弱随机控制的梯度型动力学。尽管目标函数在漂移项上通常是非凸的,但一致凸的运行代价会诱导出相应路径律的严格凸泛函,从而产生自然的Bregman几何。我们构造了相关的隐式近端格式,并证明其适定性和收敛到连续时间镜像流。当镜像代价与运行代价一致且律依赖势为线性时,该流在中心熵坐标中变为仿射。这产生了所得牛顿流的变分构造、一致稳定性估计以及到优化器的指数收敛。对于一般的一致凸镜像代价和律的非线性凸势,我们利用二次BSDE的BMO估计在全局学习时间内构造自然梯度流,并建立定量收敛性和一阶离散误差。我们进一步分析相互作用的有限粒子近似,证明到有限粒子最优值的指数收敛以及到平均场最优值的O(1/N)归一化差异。在马尔可夫情形下,反射耦合在Rd上产生反馈控制的一致空间正则性和指数收敛,包括非二次状态依赖哈密顿量的情形。

英文摘要

We study gradient-type dynamics for weak stochastic control through the laws of controlled diffusions on path space. Although the objective is generally nonconvex in the drift, a uniformly convex running cost induces a strictly convex functional of the corresponding path law and hence a natural Bregman geometry. We construct the associated implicit proximal scheme and prove its well-posedness and convergence to a continuous-time mirror flow. When the mirror cost coincides with the running cost and the law-dependent potential is linear, the flow becomes affine in centered entropy coordinates. This yields a variational construction of the resulting Newton flow, uniform stability estimates, and exponential convergence to the optimizer. For a general uniformly convex mirror cost and a nonlinear convex potential of the law, we construct the natural-gradient flow globally in learning time using BMO estimates for quadratic BSDEs, and establish quantitative convergence and a first-order discretization error. We further analyze interacting finite-particle approximations, proving exponential convergence to the finite-particle optimum and an O(1/N) normalized discrepancy from the mean-field optimum. In the Markovian setting, reflection coupling yields uniform spatial regularity and exponential convergence of the feedback controls on Rd, including for nonquadratic state-dependent Hamiltonians.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑