EgoExo-WM: 利用外部视频解锁自我世界模型
EgoExo-WM: Unlocking Exo Video for Ego World Models
浏览论文内容
中文总结 AI 辅助
提出通过从外部视频提取结构化身体姿态并利用人体运动学先验将其转换为自我视频,从而利用丰富的野外外部数据训练自我世界模型,显著提升预测质量和下游规划性能。
中文摘要 AI 辅助
自我中心世界模型为智能体预测和规划提供了有前景的方向,但其性能受限于自我中心训练数据的有限性以及人类物理动作的固有部分可观测性。相比之下,外部中心视频丰富且能很好地揭示身体姿态,但缺乏与智能体动作空间的直接对齐,且不是自我中心的。我们提出一种方法,通过从外部中心视频中提取结构化身体姿态作为动作表示,并基于人体运动学先验将外部中心视频转换为自我中心视频,从而弥合这一差距。这一过程使得将野外外部中心数据整合到自我中心世界模型训练中成为可能。我们表明,使用转换后的数据训练全身动作条件自我中心世界模型显著提高了预测质量和下游规划性能,其中我们推断实现视觉目标状态所需的身体姿态序列。我们的方法为利用任意野外视频构建强大的自我中心世界模型铺平了道路,进一步推动了机器人规划和增强现实指导等应用。
英文摘要
Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans' physical actions. In contrast, exocentric video is abundant and reveals body poses well, but lacks direct alignment with an agent's action space -- and is not egocentric. We propose a method to bridge this gap by extracting structured body pose from exocentric video as a representation of action and transforming the exocentric video to egocentric video, informed by a human kinematics prior. This process unlocks the integration of in-the-wild exocentric data for egocentric world model training. We show that training whole-body action-conditioned egocentric world models with our converted data significantly improves both prediction quality and downstream planning performance, where we infer the sequence of body poses needed to achieve a visual goal state. Our approach paves the way to enlist arbitrary in-the-wild videos for building powerful egocentric world models, furthering applications in robot planning and augmented-reality guidance.
发表机构
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。