AI 中文总结
该研究针对人形机器人长程移动操作的复杂序列任务,提出分层基于模型的强化学习框架LUCID,经模拟多物体重排实验验证,其任务成功率优于现有基线方法。
AI 中文摘要
长程人形机器人移动操作需要组合多种全身技能以及可靠的高层决策。现有方法通常将预训练技能与脚本规划器、有限状态机或任务特定的无模型策略进行协调,限制了其处理复杂任务序列的能力。为解决这一局限,我们提出LUCID,这是一种分层的基于模型的强化学习框架,通过学习到的动力学模型的想象滚动,在可复用技能上进行规划。LUCID首先通过对抗式模仿训练结构化的潜在条件低层策略,随后冻结该策略,同时联合学习高层策略和宏观动力学世界模型。该世界模型能够预测由潜在决策引发的时间扩展状态转移,支持通过想象滚动进行高层策略优化。我们在多种模拟多物体重排场景中评估了该框架,实验结果表明,与现有基线方法相比,LUCID提升了全任务成功率和部分完成率,证明了其在复杂序列移动操作任务中的有效性。
英文摘要
Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their ability to handle complex task sequences. To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model. LUCID first trains a structured latent-conditioned low-level policy via adversarial imitation and then freezes it while jointly learning a high-level policy and macro-dynamics world model. The world model predicts the temporally extended state transitions induced by latent decisions, enabling high-level policy optimization through imagined rollouts. We evaluate our framework across various simulated multi-object rearrangement scenarios. Experimental results show that LUCID improves the full-task success and partial-completion rates compared to prior baseline methods, demonstrating its effectiveness in complex sequential loco-manipulation tasks.