arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14548math.OC

SCMO:无穷维空间上概率测度优化的随机控制方法

SCMO: Stochastic Control for Optimization over Probability Measures on Infinite-Dimensional Spaces

Hang Cheung, Jinniao Qiu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SCMO,一种基于熵正则化随机控制的无梯度粒子优化方法,用于在无穷维空间上优化概率测度,能处理非凸非光滑目标并恢复多模态分布。

中文摘要 AI 辅助

我们研究在可分希尔伯特空间上的概率测度上,对可能非凸且非光滑的泛函进行仅目标优化的方法,允许优化器本质上为非狄拉克测度。我们引入了SCMO(随机控制测度优化器),这是一种源自熵正则化随机控制的无梯度粒子方法。经过有限粒子和伽辽金近似后,Cole-Hopf变换将最优反馈表示为吉布斯加权的终端位移。SCMO通过采样上下文云来近似此反馈,每次用一个候选样本替换一个粒子,对所得经验测度进行评分,并应用指数重加权。SCMO对候选提议和粒子更新使用独立的协方差,当相等时称为匹配,否则称为非匹配;我们的分析涵盖这两种设置。在匹配情况下,我们建立了无PDE的定性收敛性以及一个先投影后定量的界,该界将粒子、伽辽金投影和熵正则化误差分开。对于具有非匹配协方差的实际多上下文实现,我们证明了有限候选一致性,并表明非匹配协方差的有符号、曲率相关效应可以降低近似误差的上界。最后,在函数空间、轨迹分布和接触丰富的操作问题上的实验表明,SCMO能够处理非光滑和非凸目标,逃离次优局部盆地,并恢复规定的多模态、非狄拉克分布结构。复现我们实验结果的代码可在以下网址获取:此https URL。

英文摘要

We study objective-only optimization of possibly nonconvex and nonsmooth functionals over probability measures on a separable Hilbert space, allowing the optimizer to be intrinsically non-Dirac. We introduce SCMO (Stochastic Control Measure Optimizer), a gradient-free particle method derived from entropy regularized stochastic control. After finite-particle and Galerkin approximations, a Cole--Hopf transform represents the optimal feedback as a Gibbs-weighted terminal displacement. SCMO approximates this feedback by sampling context clouds, replacing one particle at a time with candidate draws, scoring the resulting empirical measures, and applying exponential reweighting. SCMO uses separate covariances for candidate proposals and particle updates, termed matched when equal and nonmatched otherwise; our analysis covers both settings. In the matched case, we establish PDE-free qualitative convergence and a projection-first quantitative bound separating particle, Galerkin-projection, and entropy-regularization errors. For the practical multi-context implementation with nonmatched covariance, we prove finite-candidate consistency and show that the signed, curvature-dependent effect of nonmatched covariance can reduce the resulting upper bound on the approximation error. Finally, experiments on function-space, trajectory-law, and contact-rich manipulation problems show that SCMO handles nonsmooth and nonconvex objectives, escapes suboptimal local basins, and recovers prescribed multimodal, non-Dirac law structure. Code for reproducing our experimental results is available at https://github.com/HenryCHEUNG7373/SCMO.

发表机构

  • University of Calgary(卡尔加里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑