发表机构
NVIDIA; Brown University; Columbia University; Harvard University(英伟达; 布朗大学; 哥伦比亚大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Hydra-0通用世界模型,以动作流为条件实现跨实体、任务等的通用建模控制,在RoboLab基准获高相关性,还涌现出逆模式,可助力机器人控制。
AI 中文摘要
我们提出Hydra-0,这是一种以动作流为条件的通用世界模型,它将机器人动作表示为像素运动。这种共享视觉界面通过学习跨不同实体、任务、环境和视频生成主干的动作后果,实现通用世界建模与控制。我们的最优配置相较于以动作为条件的基线,机器人运动误差降低90.4%,物体运动误差降低60.2%,同时支持零样本组合与数据高效适配。在RoboLab基准上,Hydra-0在重放成功率与参考成功率之间达到皮尔逊相关系数r=0.96。最后,我们揭示了该界面的一种涌现逆模式:一种世界动作模型,可从人类演示传递的期望物体流中预测兼容的机器人运动。训练好的动作头将所得潜在特征映射为可执行动作,无需特定任务的专家机器人演示。这些结果共同证明,动作流作为共享控制界面,在连接异构训练数据、开环策略评估与机器人控制方面具有潜力。
英文摘要
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.
CommentsProject page: https://nvidia-isaac.github.io/video_to_data/hydra-0/