发表机构
SJTU(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UniMoCa是统一运动与相机控制的视觉代理框架,提出MCVP表征并构建MCVP-Video数据集,在Wan2.2 I2V上实现视频生成多方面性能提升且复杂度低
AI 中文摘要
控制人体运动和相机移动对于忠实的人类导向视频生成至关重要,但在存在大幅肢体动作、遮挡和动态相机的多人场景中仍具挑战性。现有流程通常依赖骨骼图、姿态图或渲染人体表征等视觉运动序列进行运动控制,同时使用相机嵌入进行相机控制。这种异构控制界面迫使视频生成模型协调像素对齐的视觉线索与非视觉几何嵌入,导致运动-相机归因困难且对相机估计误差敏感。我们提出UniMoCa,一种在视觉空间中统一运动与相机控制的表征驱动框架。UniMoCa的核心是运动-相机视觉代理(MCVP),这是一种可共享的新型表征,将从驱动视频中提取的3D人体运动和相机轨迹转换为身份中立的视觉代理。MCVP在恢复的相机轨迹下渲染时间对齐的人体几何,并添加显式相机轨迹标记,以可区分的视觉线索取代异构视觉-参数控制。由于两个控制因子在同一视觉空间中表示,它们变得相互兼容而非异构,支持视频生成过程中的一致联合推理与编辑。我们还整理了MCVP-Video数据集,涵盖复杂动作、多人交互和多样相机轨迹。基于Wan2.2 I2V的实验表明,UniMoCa在人体运动控制、相机控制、时间一致性和相机感知鲁棒性方面取得显著提升,且额外复杂度极低。更多详情见我们的项目页面:this https URL
英文摘要
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences, such as skeleton maps, pose maps, or rendered body representations, for motion control, while using camera embeddings for camera control. Such heterogeneous control interfaces force video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings, making motion-camera attribution difficult and sensitive to camera estimation errors. We propose \textbf{UniMoCa}, a representation-driven framework that unifies motion and camera controls in visual space. At the core of UniMoCa is \textbf{Motion-Camera Visual Proxy} (\textbf{MCVP}), a mutually-sharable novel representation that converts 3D human motion and camera trajectories extracted from driving videos into an identity-neutral visual proxy. MCVP renders temporally aligned human geometry under the recovered camera trajectory and augments it with explicit camera trajectory markers, replacing heterogeneous visual-parametric controls with distinguishable visual cues. As both control factors are represented in the same visual space, they become mutually compatible rather than heterogeneous, enabling consistent joint reasoning and editing during video generation. We further curate a \textbf{MCVP-Video} dataset covering complex actions, multi-person interactions, and diverse camera trajectories. Experiments based on the Wan2.2 I2V show that UniMoCa achieves substantial gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity. More details are shown in our Project page: https://tanliming-daniel.github.io/UniMoCa/.