发表机构
Kuaishou Technology; Peking University(快手科技; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有运动迁移依赖固定结构对应关系的局限,提出两阶段框架,结合自研数据集与基准,实现跨类别运动迁移,在运动保真度和目标保留上达当前最优。
AI 中文摘要
视频运动迁移旨在利用参考视频的动力学为目标对象赋予动态效果。现有方法大多依赖固定的结构对应关系,当参考与目标对象在形态、关节结构或变形机制上差异显著时,该对应关系会失效。我们提出“超越形态的运动”这一思路,旨在通过保留在不同目标形态间仍有意义的动力学,实现超越固定结构对应关系的运动迁移。为实现这一目标,我们设计了两阶段框架:第一阶段学习互补的多粒度抽象运动视图,并利用这些视图引导跨类别视频对,保留不同形态间可迁移的动力学;第二阶段将该监督信号内化到直接基于参考视频条件的生成中,推理时无需显式提取运动。我们还推出了用于训练和评估图像、文本条件下运动迁移的OpenVMT-Dataset与OpenVMT-Bench,覆盖同类别、近类别、远类别差距,计划录用后发布。大量实验表明,该方法在运动保真度和目标保留方面达到了当前最优水平。项目页面:this https URL
英文摘要
Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations largely rely on fixed structural correspondence, which becomes ill-defined when reference and target objects differ substantially in morphology, articulation, or deformation mechanisms. We introduce Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies. To realize this, we propose a two-stage framework. Stage~I learns complementary multi-granularity abstract motion views and uses them to bootstrap cross-category video pairs that preserve transferable dynamics across diverse morphologies. Stage~II internalizes this supervision into direct reference-video-conditioned generation, removing the need for explicit motion extraction at inference. We further introduce OpenVMT-Dataset and OpenVMT-Bench for training and evaluating image- and text-conditioned motion transfer across Same, Near, and Far category gaps. Extensive experiments demonstrate state-of-the-art motion fidelity and target preservation. Project page: https://miniz233.github.io/MotionBeyondMorphology/