arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于样本高效模型预测控制的混合反馈采样

Hybrid Feedback Sampling for Sample-Efficient Model Predictive Control

Chaoyi Pan, Zeji Yi, John Zhang, Zachary Manchester, Guannan Qu, Guanya Shi

arXiv 2608.19443首次发表:更新:

发表机构

Carnegie Mellon University; Massachusetts Institute of Technology(卡内基梅隆大学; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出混合反馈采样MPC(FS-MPC)算法,平衡局部与全局搜索,解决采样MPC在高维开环不稳定系统中的样本效率与数值不稳定问题,在真实机器人控制任务中表现优于现有方法。

AI 中文摘要

基于采样的模型预测控制(MPC)因具备并行性与灵活性,已广泛应用于真实机器人系统的控制。但对于高维、开环不稳定的动力系统,用于优化控制序列所需的样本数量会随预测时域呈指数增长,导致样本效率低下且数值不稳定。本文研究了基于采样的MPC中射击法的不稳定性,表明最优采样提议分布可通过优化反馈策略采样实现,将该算法命名为反馈采样MPC(FS-MPC)。FS-MPC采用混合采样设计,基于系统稳定性与可用计算预算平衡局部与全局搜索。理论分析显示,该混合采样方法比标准MPPI收敛更快,且比标准反馈采样的最优性更好。在人形机器人运动操作、灵巧操作等多种接触丰富的控制任务中,FS-MPC成功解决了标准采样方法难以处理的动态不稳定任务,且严格优于单独的反馈策略;最终在真实世界的人形机器人运动与操作任务中验证了该方法的有效性。

英文摘要

Thanks to its parallelizability and flexibility, sampling-based Model Predictive Control (MPC) has become widely popular for controlling real-world robotic systems. However, for high-dimensional and open-loop unstable dynamical systems, the required number of samples to improve the control sequence will grow exponentially with the horizon, leading to poor sample efficiency and numerical instability. This paper investigates the instability of shooting methods in sampling-based MPC and shows that the optimal sampling proposal distribution can be realized by sampling with an optimized feedback policy. We refer to this algorithm as Feedback Sampling MPC (FS-MPC). FS-MPC involves a hybrid sampling design which balances local and global search based on the system stability and the available computation budget. Our theoretical analysis shows that our hybrid sampling approach achieves faster convergence than standard MPPI and better optimality than standard feedback sampling. Empirically, in diverse contact-rich control tasks like humanoid loco-manipulation and dexterous manipulation, we show that FS-MPC successfully tackles dynamically unstable tasks where standard sample-based approaches struggle, and strictly outperforms feedback policies alone. Finally, we validate our method on humanoid robot locomotion and manipulation tasks in the real world.

Comments15 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑