发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReCo通过响应塑形和闭环响应模型使MPC联合规划运动与手臂动作,在仿真中降低位置和方向误差28.7%和27.4%,并实现真实机载连续腿部操作。
AI 中文摘要
连续腿部操作要求在基座持续行走的同时实现精确的末端执行器跟踪。将强化学习(RL)与模型预测控制(MPC)相结合适用于该任务:学习到的策略提供稳健的运动,而MPC协调基座和手臂以补偿跟踪误差。然而,MPC只能补偿其能够预测的基座运动,且学习策略的命令响应会随步态相位、接触和负载而变化。我们提出ReCo,一个将响应一致运动与策略感知MPC相结合的腿部操作框架。响应塑形训练策略在随机化动力学下对命令做出一致且可重复的响应。随后,一个辨识出的闭环响应模型使MPC能够联合规划运动命令和手臂运动。在仿真基准上,ReCo将位置和方向均方根误差(RMSE)分别相对于各指标的最佳基线降低了28.7%和27.4%。真实世界实验展示了具有协调基座和手臂运动的机载连续腿部操作。
英文摘要
Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking. Combining reinforcement learning (RL) with model predictive control (MPC) suits this task: the learned policy provides robust locomotion, while MPC coordinates the base and arm to compensate for tracking errors. However, MPC can compensate only for base motion that it can predict, and a learned policy's command response varies with gait phase, contact, and payload. We present ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics. An identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion. On the simulation benchmark, ReCo reduces position and orientation root-mean-square error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.