World2Motion:将视频世界模型转化为3D人体运动生成器
World2Motion: Turning Video World Models into 3D Human Motion Generators
浏览论文内容
中文总结 AI 辅助
World2Motion将视频世界模型Cosmos 3改造为单阶段3D运动生成器,通过构建混合训练数据和移位解耦噪声调度,实现从单图与文本生成场景感知运动,提升对齐与交互并加速推理3.3倍。
中文摘要 AI 辅助
我们提出了World2Motion,这是一个从单张图像和文本提示生成场景感知的3D人体运动及对应视频的框架。现有的3D运动生成器从运动数据集中学习,但其泛化能力受限于环境覆盖范围有限。相比之下,诸如Cosmos 3之类的视频世界模型提供了更广泛的环境先验,但并非为全身运动生成而设计;从它们生成的视频中恢复运动需要昂贵的两阶段推理。为解决这些问题,我们将Cosmos 3转变为单阶段3D运动生成器。这一改造面临两个挑战:配对视频-运动数据的稀缺性以及生成运动的时间不稳定性。首先,我们构建了一个训练数据集,结合了合成视频-运动对与配对估计3D运动的真实视频。其次,我们提出了一种移位解耦噪声调度,通过共享的去噪进度为视频和运动分配不同的噪声水平。这种设计适应了两种模态不同的去噪需求,减少了运动抖动。在多源交互基准上的实验表明,与所评估的3D运动生成器相比,World2Motion具有更好的运动-文本对齐和场景交互。它还匹配了两阶段基线的交互成功率,同时实现了约3.3倍的推理速度提升。
英文摘要
We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3$\times$ faster inference. Our project page is available at https://fyantu.github.io/World2Motion/.
发表机构
- Japan Advanced Institute of Science and Technology(日本先端科学技术大学院大学)
- Institute of Science Tokyo(东京科学大学)
- The University of Tokyo Alaya Lab(东京大学Alaya实验室)
机构由 AI 辅助整理,请以论文原文为准。