AI 中文总结
MotionPhys是基于光流轨迹物理一致性的轻量可解释框架,可检测AI生成视频的物理运动不一致性,且对不同视频生成器泛化性良好。
AI 中文摘要
现代AI视频生成模型可生成视觉保真度高、时间过渡看似平滑的视频,但视觉真实度未必意味着物理运动一致性。现有生成模型主要优化像素或隐空间的分布匹配,未明确施加惯性、连续力、轨迹几何等现实世界约束。我们的实验表明,AI生成视频在短帧序列上仍具视觉合理性,但在完整物体动作过程中无法保持物理运动一致性,导致运动轨迹存在系统性统计差异。基于该观察,我们提出MotionPhys,这是一个轻量且可解释的框架,将稀疏运动轨迹视为物理证据,而非依赖外观伪影或生成器特定痕迹。通过建模轨迹在多个时间尺度上的几何演化,MotionPhys可揭示常规视觉线索难以捕捉的细微运动不一致性,并将其转化为紧凑表示以实现高效检测。在多个数据集上的实验显示,MotionPhys可有效检测生成视频中的物理不一致性,且对不同视频生成器具有良好泛化能力。
英文摘要
Modern AI video generation models can produce videos with high visual fidelity and seemingly smooth temporal transitions. However, visual realism does not necessarily imply physical motion consistency. Existing generative models mainly optimize distribution matching in pixel or latent spaces, without explicitly enforcing real-world constraints such as inertia, continuous forces, and trajectory geometry. Our experiments show that AI-generated videos remain visually plausible over short sequences of consecutive frames, yet fail to preserve physical motion consistency throughout a complete object action, resulting in systematic statistical discrepancies in their motion trajectories. Based on this observation, we introduce MotionPhys, a lightweight and interpretable framework that treats sparse motion trajectories as physical evidence rather than relying on appearance artifacts or generator-specific traces. By modeling the geometric evolution of trajectories across multiple temporal scales, MotionPhys reveals subtle motion inconsistencies that are difficult to capture with conventional visual cues and transforms them into a compact representation for efficient detection. Experiments on multiple datasets show that MotionPhys can effectively detect physical inconsistencies in generated videos and generalizes well across different video generators.