arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12347cs.RO

DWMP:利用双世界模型实现人形机器人跨越障碍

DWMP: Leveraging Dual World Models for Humanoid Obstacle Traversal

Rongjun Jin, Jianming Ma, Yue Gao

首次发表
浏览论文内容

中文总结 AI 辅助

针对人形机器人穿越障碍时多模态观测特性差异问题,提出双世界模型策略,分别用Koopman和RSSM建模动力学与视觉,融合表示生成动作,在仿真和真机上提升性能。

中文摘要 AI 辅助

人形机器人必须利用机载本体感觉和视觉观察穿越杂乱无章的障碍区域,然而现有方法通常处理多模态观测时并未明确考虑它们的不同特性:本体感觉观测维度低,但受高度非线性的机器人动力学支配;而自我中心视觉观测维度高、噪声大且冗余。我们提出DWMP(双世界模型策略),该框架为智能体提供分离但互补的世界模型表示,用于人形机器人跨越障碍。基于Koopman的动力学世界模型将本体感觉观测提升到潜在空间,在该空间中其时间演化近似线性,使得动力学特征更易于智能体学习。基于RSSM的视觉世界模型将自我中心深度观测压缩为紧凑的随机状态,同时保留与障碍相关的几何信息。学生策略接收融合的潜在表示以生成动作,将线性化的本体感觉动力学与压缩的视觉感知相结合。在仿真和Unitree G1人形机器人上的实验表明,DWMP在跨越障碍性能上优于基线方法,并支持在随机障碍布局下的现实世界部署。

英文摘要

Humanoid robots must traverse cluttered obstacle fields using onboard proprioceptive and visual observations, yet existing methods usually process multimodal observations without explicitly considering their different characteristics: proprioceptive observations are low-dimensional but governed by highly nonlinear robot dynamics, while egocentric visual observations are high-dimensional, noisy, and redundant. We propose DWMP (Dual World Model Policy), a framework that provides the actor with separate but complementary world-model representations for humanoid obstacle traversal. A Koopman-based dynamics world model lifts proprioceptive observations into a latent space where their temporal evolution is approximately linear, making the dynamics features easier for the actor to learn from. An RSSM-based visual world model compresses egocentric depth observations into compact stochastic states while preserving obstacle-related geometry. The student policy receives the fused latent representation for action generation, combining linearized proprioceptive dynamics with compressed visual perception. Experiments in simulation and on a Unitree G1 humanoid robot show that DWMP improves obstacle traversal performance over baselines and supports real-world deployment under randomized obstacle layouts.

↑