arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于稀疏离线到在线强化学习从SMPC演示中学习移动操作

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Brüdigam

arXiv 2608.12063首次发表:更新:

发表机构

RAI Institute; Technical University of Munich; ETH Zurich(RAI研究所; 慕尼黑工业大学; 苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用仿真中SMPC生成的离线数据,结合稀疏奖励训练离策略RL智能体,整合动态稳定性控制器,实现机器人移动操作技能的仿真到真实迁移,性能超越原最优控制方法。

AI 中文摘要

将移动与操作功能整合对机器人自主至关重要,但将标准强化学习(RL)扩展至复杂任务时,会因密集奖励塑造的缓慢手动过程而严重受阻。为规避这一局限,我们完全在仿真中利用基于样本的模型预测控制(SMPC)作为自动、可快速调优的专家,生成大规模离线数据集。由于该数据解决了基础探索问题,我们可仅使用稀疏任务奖励训练离策略RL智能体,大幅缩短学习新技能的时间并消除手动调优需求。将该高层智能体与低层动态稳定性控制器整合,可产生更优行为,严格契合真实任务目标,最终使所学策略超越原最优控制教师。我们通过在不同形态机器人上成功部署复杂移动操作技能,验证了此仿真到真实框架的鲁棒性,这些机器人包括配备机械臂的Spot四足机器人与G1人形机器人。

英文摘要

Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC) entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets. Because this data solves the fundamental exploration problem, we can train an off-policy RL agent using purely sparse task rewards, drastically reducing the time required to learn new skills and eliminating the need for manual tuning. Integrating this high-level agent with a low-level dynamic stability controller yields more optimal behaviors that strictly align with true task objectives, ultimately allowing the learned policies to surpass the original optimal control teacher. We validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies, including an arm-equipped Spot quadruped and a G1 humanoid.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑