基于世界模型增强的人形机器人在支点受限地形上的视觉 locomotion
World-Model-Augmented Visual Locomotion for Humanoids on Foothold-Constrained Terrain
浏览论文内容
中文总结 AI 辅助
该研究提出WM-LOCO方法,通过联合训练循环世界模型与PPO策略,使人形机器人在支点受限地形上的平均成功率达93.3%,优于基线方法。
中文摘要 AI 辅助
支点受限地形的特征是可行足部接触点稀疏、不连续或几何受限,如踏脚石、间隙及狭窄楼梯踏步。在此类地形上,一步失误往往几乎无补救空间,因此仅基于即时可见地形制定足部放置决策的策略易失败。本文探究学习近未来观测与奖励的预测摘要是否能提供此类场景所需的前瞻性信息。我们提出世界模型增强的视觉 locomotion(WM-LOCO),其联合训练循环世界模型与PPO策略。在本体感知和单机载深度图像的条件下,世界模型生成预测循环特征以指导策略,无需明确的支点标签。在仿真中,WM-LOCO在间隙和踏脚石任务上的表现优于完全失败的匹配基线,在楼梯任务上与基线成功率相当,同时提升步长效率并降低骨盆加速度。我们将同一策略部署到实体Unitree G1人形机器人上,利用机载本体感知和单深度流,其在三类地形上的平均成功率达93.3%。
英文摘要
Foothold-constrained terrain is characterized by sparse, discontinuous, or geometrically restricted feasible foot contacts, as encountered on stepping stones, across gaps, and on narrow stair treads. On such terrain, a single misstep often leaves little room to recover, so policies that base foot-placement decisions primarily on the immediately visible terrain are prone to failure. We ask whether a learned predictive summary of near-future observations and rewards can provide the anticipatory information required in such settings. We present World-Model-Augmented Visual Locomotion (WM-LOCO), which jointly trains a recurrent world model and a PPO policy. Conditioned on proprioception and a single onboard depth image, the world model produces a predictive recurrent feature that guides the policy, without explicit foothold labels. In simulation, WM-LOCO succeeds on gaps and stepping stones where a matched baseline fails completely, and matches the baseline's success rate on stairs while improving stride efficiency and reducing pelvis acceleration. We deploy the same policy onboard a physical Unitree G1 humanoid using onboard proprioception and a single depth stream; it traverses all three terrain classes with an average success rate of 93.3%.
发表机构
- D-Robotics(地平线机器人)
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- Soochow University(苏州大学)
机构由 AI 辅助整理,请以论文原文为准。