发表机构
NTU; PKU; BAAI; HKUST(GZ)(南洋理工大学; 北京大学; 北京智源人工智能研究院; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 ω-0 模型,针对类人机器人协同移动操作问题,结合潜在视觉预见与扩散动作生成,在 11 项家庭任务上优于多个基线,还构建了 ω-HOME 数据集。
AI 中文摘要
类人机器人的 household 任务通常需要协同移动操作,即机器人必须将移动、调整姿态、保持平衡和操作对象整合为单一协调行为。然而,现有类人机器人策略通常将移动与操作解耦,而近期的世界动作模型要么以机械臂为中心,要么以视频为中心。本文提出 ω-0,一种用于真实世界类人机器人协同移动操作的潜在预测全身世界动作模型。给定语言指令、当前视觉观测和机器人本体感受状态,ω-0 直接预测控制器兼容的全身动作潜在变量,用于真实机器人执行。ω-0 不重构未来视频,而是学习紧凑的未来观测嵌入作为轻量级预测目标,将潜在视觉预见与基于扩散的全身动作生成相结合。该模型支持以自我为中心的 RGB、以异中心的 RGB 和以异中心的深度输入,并利用基于控制器的仿真回放,将人类/公共视觉运动先验转化为机器人可执行的动作潜在变量。我们进一步收集 ω-HOME,这是一个时长超 40 小时的真实世界家庭类人机器人数据集,包含同步多视图观测、全身 SMPL 运动、机器人状态和动作潜在变量。针对 11 项家庭任务的真实世界实验表明,单一 ω-0 模型可生成流畅的移动中操作行为,且始终优于代表性的模仿学习、VLA、类人机器人和 WAM 基线。
英文摘要
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present $ω$-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, $ω$-0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, $ω$-0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect $ω$-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single $ω$-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.