发表机构
Yonsei University; GIST(延世大学; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多人物图像动画中身份-运动绑定和遮挡难题,提出WeLike2Party框架,采用上下文视频条件化、参考非对称RoPE条件化和身份绑定监督,并构建MotionTwin数据集与基准,实验验证其优于现有方法。
AI 中文摘要
人类图像动画旨在将驱动视频中的运动迁移到参考图像中的主体上。尽管视频生成领域取得了显著进展,但实现多个交互主体的高保真动画仍然是一个挑战。许多现有方法依赖于显式运动表示,如2D骨架或参数化身体网格,并且在人物间遮挡情况下难以保持身份-运动绑定。为解决这一局限,我们提出了WeLike2Party,一个基于直接上下文视频条件化的多人物动画框架,在推理时无需显式姿态或网格提取。我们进一步引入了参考非对称RoPE条件化(Reference Asymmetric RoPE Conditioning)以保留细粒度外观细节,以及身份绑定监督(Identity Binding Supervision)以将每个参考身份与其预期运动轨迹关联。为支持跨身份训练,我们构建了MotionTwin,一个大规模合成数据集,包含14.4K个共享主体和相机运动的跨身份视频对,总计84.3小时的逼真视频。我们还提出了MotionTwin-Bench,一个专门设计用于评估主体级视觉保真度和身份-运动绑定的跨身份基准。在MotionTwin-Bench和真实世界视频上的大量实验表明,WeLike2Party在主体级视觉保真度、身份-运动绑定和整体感知质量方面优于近期最先进的方法,特别是在具有大量遮挡的多人物交互场景中。
英文摘要
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.