发表机构
Cardiff University; University of Science and Technology of China; East China Normal University(卡迪夫大学; 中国科学技术大学; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MotionSpec框架,通过谱轨迹一致性(STC)和局部流一致性(LFC)监督,提升视频生成中的运动一致性、时间连贯性与合理性,同时保持视觉保真度。
AI 中文摘要
近年来,文本到视频生成的进展实现了高保真度的视觉合成,但逼真的运动仍然具有挑战性。生成的视频可能表现出时间不连续性、动作进展不一致以及复杂运动过程中的结构变形。即使单个帧看起来逼真,底层运动也可能以不一致或不合理的方式演变。标准生成目标提供的运动特定监督有限,导致运动演变约束不足。在本文中,我们提出了MotionSpec,一个以谱轨迹一致性(STC)为中心的运动监督框架。STC构建密集的锚点相对运动轨迹,并通过时间傅里叶变换将其转换为运动谱体积。通过对齐预测轨迹和目标轨迹的谱幅度和相位,STC约束了跨时间频率的运动强度以及运动的时间组织。为了补充这种轨迹级监督,我们引入了局部流一致性(LFC),它对齐预测视频和目标视频之间的连续帧光流,以稳定局部运动过渡。实验表明,MotionSpec在保持视觉保真度的同时,持续提高了运动一致性、时间连贯性和合理性。
英文摘要
Recent advances in text-to-video generation have enabled high-fidelity visual synthesis, yet realistic motion remains challenging. Generated videos may exhibit temporal discontinuities, inconsistent action progression, and structural distortions during complex movements. Even when individual frames appear realistic, the underlying motion may evolve in inconsistent or implausible ways. Standard generative objectives provide limited motion-specific supervision, leaving motion evolution insufficiently constrained. In this paper, we propose MotionSpec, a motion supervision framework centered on Spectral Trajectory Consistency (STC). STC constructs dense anchor-relative motion trajectories and transforms them into motion spectral volumes via a temporal Fourier transform. By aligning the spectral amplitude and phase of predicted and target trajectories, STC constrains both motion strength across temporal frequencies and the temporal organization of motion. To complement this trajectory-level supervision, we introduce Local Flow Consistency (LFC), which aligns consecutive-frame optical flow between predicted and target videos to stabilize local motion transitions. Experiments demonstrate that MotionSpec consistently improves motion consistency, temporal coherence, and plausibility while preserving visual fidelity.