arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10332eess.SYcs.LGcs.SY

可微预测控制的拓扑可行性保证

Topological Feasibility Guarantees for Differentiable Predictive Control

Guangyu Wu, Ján Drgoňa

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对可微预测控制(DPC)缺乏严格可行性保证的问题,通过拓扑分析建立其确定性可行性保证,提出结合控制障碍函数(CBFs)的自监督离线策略学习策略,仿真验证其约束违反量随样本规模增大趋近于零,为学习型控制的可行性保证提供新视角。

中文摘要 AI 辅助

可微预测控制(Differentiable Predictive Control, DPC)是一种用于近似显式模型预测控制(Model Predictive Control, MPC)策略的自监督学习方法,相较于基于在线优化的MPC具有显著的计算优势。然而,作为安全控制核心要求的可行性保证,目前仅能通过概率方式或在线安全滤波器提供。离线策略优化缺乏严格的可行性保证仍是一个开放性问题。本文通过对诱导可达安全集进行新颖的拓扑分析,为DPC建立了确定性可行性保证,无需在线安全滤波器。利用DPC固有的基于模型的特性(可微系统动力学直接嵌入计算图中),我们从拓扑和几何角度分析了学习到的控制策略及对应系统状态的属性。受理论分析启发,我们提出了一种新颖的自监督离线策略学习策略,该策略利用带有控制障碍函数(Control Barrier Functions, CBFs)的代理损失。关键在于,这些属性不仅能显著改进策略训练,还能基于有限数量的训练样本推导出严格的确定性可行性保证。大量闭环仿真验证了我们的理论发现,表明随着训练样本规模的增加,经验约束违反量单调递减至零。最终,本研究表明DPC策略优化能产生形式化安全证书,而传统黑盒方法(如强化学习(Reinforcement Learning, RL)或基于监督学习的近似MPC)在结构上无法获得此类证书,从而为基于学习的控制中的可行性保证提供了新视角。

英文摘要

Differentiable predictive control (DPC), a self-supervised learning approach for approximating explicit model predictive control (MPC) policies, offers significant computational advantages over online optimization-based MPC. However, feasibility guarantees, a core requirement for safe control, are currently provided either probabilistically or via online safety filters. The lack of rigorous feasibility guarantees for offline policy optimization remains an open problem. This paper establishes deterministic feasibility guarantees for DPC using a novel topological analysis of the induced reachable safe set, without requiring online safety filters. By exploiting the inherent model-based nature of DPC, in which differentiable system dynamics are embedded directly into the computational graph, we analyze the properties of the learned control policies and the corresponding system states from topological and geometric perspectives. Inspired by our theoretical analysis, we propose a novel self-supervised offline policy learning strategy that utilizes a proxy loss with Control Barrier Functions (CBFs). Crucially, these properties not only significantly improve policy training but also enable the derivation of strict, deterministic feasibility guarantees from a finite number of training samples. Extensive closed-loop simulations validate our theoretical findings, demonstrating that the empirical constraint violations monotonically decrease to zero as the training sample size increases. Ultimately, this work illustrates that DPC policy optimization yields formal safety certificates that are structurally unattainable with conventional black-box methods, e.g., reinforcement learning (RL) or supervised learning-based approximate MPC, thereby providing a new perspective on feasibility guarantees in learning-based control.

发表机构

  • Chalmers University of Technology(查尔姆斯理工大学)
  • Nanyang Technological University(南洋理工大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑