arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据驱动的自适应采样非专家示范MPC:性能保证与样本复杂度

Data-Driven MPC with Adaptively Sampled Non-Expert Demonstrations: Performance Guarantees and Sample Complexity

Shijie Pan, Agustin Castellano, Enrique Mallada

arXiv 2609.22828首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于非专家示范轨迹的数据驱动MPC策略,通过构建Lipschitz正则化上包络提供性能证书,并给出相对最优性保证与样本复杂度,在火箭着陆任务中验证了接近最优性能。

AI 中文摘要

我们研究利用由代理有限时域模型预测控制(MPC)生成的轨迹来为动态系统构建数据驱动策略,目标是逼近无限时域最优控制策略。我们将这些轨迹解释为非专家示范,并提出一种基于记忆的非参数策略,该策略构建有限时域代理值函数的Lipschitz正则化上包络,作为所得策略值函数的显式上界,从而在运行时绕过显式优化的同时提供可计算的性能证书。我们建立了相对于无限时域最优值函数的相对最优性保证,并推导了实现指定相对误差所需的充分MPC时域条件以及样本复杂度界。在火箭着陆任务上的数值实验验证了理论预测,并展示了接近最优的性能。

英文摘要

We study data-driven policy construction for dynamical systems using trajectories generated by surrogate finite-horizon Model Predictive Control (MPC), with the goal of approximating an infinite-horizon optimal control policy. We interpret these trajectories as non-expert demonstrations and propose a memory-based, nonparametric policy that constructs a Lipschitz-regularized upper envelope of the finite- horizon surrogate value function, which serves as an explicit upper bound on the value of the resulting policy, thereby providing a computable performance certificate while bypassing explicit optimization at runtime. We establish relative optimality guarantees with respect to the infinite-horizon optimal value function and derive sufficient MPC horizon conditions, together with sample complexity bounds, for achieving a prescribed relative error. Numerical experiments on a rocket landing task validate the theoretical predictions and demonstrate near-optimal performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑