CLAP:跨具身视频世界模型是零样本物理模拟器
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
查看机构详情
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
CLAP是跨具身动作条件视频生成框架,可在人类与机器人的多样化视频上训练,能接近或超越单具身视频模型性能,开源代码与模型,为训练单具身视频世界模型提供新范式。
中文摘要 AI 辅助
现有最先进的动作条件视频模型通常仅适用于单一机器人具身,无法利用包含丰富通用物理学习信号的大量异构视频数据。为弥合这一差距,我们提出CLAP框架,该框架用于跨具身动作条件视频生成,可在涵盖人类与机器人智能体的多样化互联网规模视频上训练。CLAP基于核心洞见:通用物理定律支配时空动态,与动作主体无关。然而跨具身学习并非易事,因为动作表示在不同机器人平台间差异显著,且人类视频中通常缺失动作信息。CLAP通过以下核心贡献解决这一根本挑战:首先,CLAP通过末端执行器位姿、语言指令和潜在动作协调不同的动作空间;其次,为克服各自局限性,CLAP引入基于课程的跨具身学习方案,先利用潜在动作从未标记视频数据中学习基础物理先验,再将其适配到末端执行器动作空间,实现零样本部署到现实任务。关键的是,CLAP在DROID等挑战性环境中性能接近或超越最先进的单具身视频模型,该性能优势通过少样本适配进一步增强,建立了训练单具身视频世界模型的新范式。最终,CLAP提供了迄今为止最全面的动作条件视频世界模型套件,涵盖多样化的动作条件空间(末端执行器、语言、潜在)和机器人形态(包括跨具身、DROID、Bridge、双机械臂YAM机器人及G1人形机器人)。我们开源了所有代码和模型,项目网站在此https URL。
英文摘要
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io .