arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10777cs.LGmath.OCstat.ML

线性二次随机最优控制的路径积分值匹配

Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control

  • Westlake University(西湖大学)
  • School of Engineering(工程学院)

机构由 AI 辅助整理,请以论文原文为准。

Bangyan Liao, Chenglei Yu, Yuchen Yang, Chuanrui Wang, Zhisheng Song, Peidong Liu, Tailin Wu

AI总结:

该研究针对线性二次随机最优控制中基于策略方法的高成本与不稳定性问题,提出路径积分值匹配算法,结合时间差分学习、吉尔萨诺夫定理与经验回放,在低维场景效率提升一个数量级,高维场景缓解模式坍塌,为复杂SOC问题提供可扩展解决方案。

AI中文摘要:

线性二次随机最优控制(LQ-SOC)为控制带噪声的动态系统构建了基础框架,近期在机器学习领域重新受到关注。然而,当前最先进的基于策略的方法因严重依赖全轨迹模拟,存在计算成本过高和不稳定的问题。为克服这些局限,我们通过重新审视路径积分控制(PIC),转向基于值的方法,实现范式转变。尽管标准PIC存在与基于策略的方法相同的高方差瓶颈,但我们发现,通过截断和边缘化原始路径积分公式,可推导出值函数的时间递归形式。基于此理论基础,我们提出路径积分值匹配(PI-VM)算法,具体而言,采用时间差分学习近似递归值动态,并结合吉尔萨诺夫定理与经验回放实现离策略训练。我们在各类SOC基准及采样任务上,将PI-VM与最先进的基于策略的方法进行基准测试。实验结果表明,在低维场景中,PI-VM达到与最先进方法相当的精度,效率提升一个数量级;在高维场景中,能有效缓解模式坍塌问题。因此,PI-VM为解决复杂SOC问题提供了可扩展方案。

英文摘要:

Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community. However, current state-of-the-art policy-based methods suffer from prohibitive computational costs and instability due to their heavy reliance on full-trajectory simulation. To overcome these limitations, we propose a paradigm shift toward a value-based approach by revisiting Path Integral Control (PIC). Although standard PIC suffers from the same high-variance bottleneck as policy-based methods, we discover that by truncating and marginalizing the original path integral formulation, we can derive a temporal recursive form of the value function. Building upon this theoretical foundation, we propose the Path Integral Value Matching (PI-VM) algorithm. Specifically, we employ temporal-difference learning to approximate the recursive value dynamics, and further integrate the Girsanov theorem with experience replay to enable off-policy training. We benchmark PI-VM against SOTA policy-based methods across various SOC benchmarks and sampling tasks. Empirical results demonstrate that PI-VM matches SOTA precision with an order-of-magnitude efficiency gain in low-dimensional settings, while effectively mitigating mode collapse in high-dimensional scenarios. Consequently, PI-VM offers a scalable solution for solving complex SOC problems.

补充信息

↑