MM-LMPC:基于模态特定终端设计与基于Bandit探索的多模态学习模型预测控制
MM-LMPC: Multi-Modal Learning Model Predictive Control via Mode-Specific Terminal Design and Bandit-Based Exploration
浏览论文内容
中文总结 AI 辅助
提出多模态LMPC(MM-LMPC),通过聚类轨迹为运动模式并构建模态特定终端设计,结合LCB元控制器选择模式,以增强探索并降低代价,理论保证稳定性与收敛性。
中文摘要 AI 辅助
学习模型预测控制(LMPC)通过利用先前的执行结果来构建MPC问题的终端约束和终端代价,从而改进迭代控制任务。尽管有效,但这种对过去轨迹的重复利用可能使LMPC对初始数据敏感。特别是,LMPC可能反复利用具有有利代价值的存储轨迹,而未能充分探索其他可能经过进一步改进后产生更低代价的替代路线模式。为解决此问题,我们提出了多模态LMPC(MM-LMPC)。所提出的框架将过去的轨迹聚类为运动模式,为每种模式构建模态特定的LMPC控制器,并使用基于LCB的元控制器在每次迭代中选择执行哪个模态特定控制器。通过两种设计将模式信息纳入终端约束和终端代价中。硬约束设计使用由每种模式相关数据构建的模态特定终端约束和终端代价。软正则化设计保留共享终端约束,同时在终端代价中添加基于隶属度的惩罚。这些设计减少了将所有轨迹汇集到单一终端记忆中所造成的偏差,同时保留了LMPC的递归可行性和稳定性结构。我们的理论分析表明,两种设计均保持了递归可行性和闭环稳定性。对于硬约束设计,我们进一步建立了模态级代价收敛性、渐近最优模态性能以及LCB规则下的对数累积遗憾界。在多路线避障任务的仿真中,MM-LMPC改善了探索性,并实现了比标准LMPC更低的代价。
英文摘要
Learning Model Predictive Control (LMPC) improves iterative control tasks by using previous executions to construct the terminal constraint and terminal cost of the MPC problem. Although effective, this reuse of past trajectories can make LMPC sensitive to the initial data. In particular, LMPC may repeatedly exploit stored trajectories with favorable cost-to-go values while insufficiently exploring alternative route patterns that could yield lower cost after further improvement. To address this issue, we propose Multi-Modal LMPC (MM-LMPC). The proposed framework clusters past trajectories into motion modes, constructs a mode-specific LMPC controller for each mode, and uses an LCB-based meta-controller to select which mode-specific controller to execute at each iteration. Mode information is incorporated into the terminal constraint and terminal cost through two designs. The hard-constrained design uses mode-specific terminal constraints and terminal costs constructed from the data associated with each mode. The soft-regularized design retains a shared terminal constraint while adding membership-based penalties to the terminal cost. These designs reduce the bias caused by pooling all trajectories into a single terminal memory while retaining the recursive feasibility and stability structure of LMPC. Our theoretical analysis shows that both designs preserve recursive feasibility and closed-loop stability. For the hard-constrained design, we further establish mode-wise cost convergence, asymptotic best-mode performance, and a logarithmic cumulative regret bound under the LCB rule. Simulations on multi-route obstacle-avoidance tasks show that MM-LMPC improves exploration and achieves lower costs than standard LMPC.
发表机构
- Graduate School of Engineering, The University of Osaka(大阪大学大学院工学研究科)
- Institute of Systems and Information Engineering, University of Tsukuba(筑波大学系统信息工程学研究所)
机构由 AI 辅助整理,请以论文原文为准。