发表机构
Amazon Prime Video; Tel Aviv University; Hebrew University of Jerusalem(亚马逊Prime Video; 特拉维夫大学; 耶路撒冷希伯来大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SemCam提出语义相机运动控制任务,通过参考视频和运动标签实现主体相对相机控制,采用低秩适应与运动调制,在109场景基准上以68.6%成功率优于基线。
AI 中文摘要
在现有视频中相对于移动主体控制相机具有挑战性:诸如保持正面视角等行为要求相机适应主体的位置和方向变化,这使得预先指定期望轨迹变得困难。现有的相机控制视频到视频方法通常依赖显式轨迹或参考运动,这些方法无法直接表达这种动态的相机-主体关系。我们引入了语义相机运动控制,这是一种新颖的视频到视频任务,其中参考视频和目标运动标签指定期望的主体相对相机行为,而无需显式目标轨迹。我们的方法SemCam学习在保留源内容的同时实现这种行为。它结合了共享基低秩适应与运动条件调制,同时背景一致性损失鼓励在参考视频和目标视频中均可见区域保持保真度。我们构建了661个覆盖八种语义相机行为的配对视频,并在一个独立的109场景基准上使用主体相对运动指标、外观度量和用户研究进行评估。SemCam实现了68.6%的语义运动成功率,而最强基线Vista4D为45.3%,同时保持了相当的主体身份保留。
英文摘要
Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to specify in advance. Existing camera-controlled video-to-video methods typically rely on explicit trajectories or reference motions, which do not directly express these dynamic camera--subject relationships. We introduce semantic camera motion control, a novel video-to-video task in which a reference video and a target motion label specify the desired subject-relative camera behavior without an explicit target trajectory. Our method, SemCam, learns to realize this behavior while preserving source content. It combines shared-basis low-rank adaptation with motion-conditioned modulation, while a background-consistency loss encourages fidelity in regions visible in both reference and target videos. We construct 661 paired videos covering eight semantic camera behaviors and evaluate on a separate 109-scene benchmark using subject-relative motion metrics, appearance measures, and a user study. SemCam achieves a semantic-motion success rate of 68.6%, compared with 45.3% for Vista4D, the strongest evaluated baseline, while maintaining comparable subject identity preservation.