发表机构
MMLab, CUHK; ByteDance Inc.; Peking University; Fudan University; CPII under InnoHK(香港中文大学多媒体实验室; 字节跳动有限公司; 北京大学; 复旦大学; 香港人工智能及机器人科技平台下的香港中文大学感知交互智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频生成中物体运动缺陷评估缺失的问题,提出MotionInsight诊断评估器,结合VidMotion数据集,从三个维度评估运动保真度,实现与人类对齐的诊断。
AI 中文摘要
尽管视频生成模型取得了快速进展,但它们仍然表现出明显的运动缺陷,通常表现为物体运动不正确。然而,大多数现有的视频质量评估侧重于美学质量或文本-视频对齐。为了解决这一差距,我们研究了以物体为中心的运动保真度评估,从物体一致性、运动连续性和物理合理性三个维度评估目标物体。为实现这一目标,我们首先引入了VidMotion,一个包含6,879个视频的诊断数据集,这些视频具有指定的运动物体和细粒度标注,包括维度得分和失败原因。我们进一步提出了MotionInsight,一个诊断评估器,将评估从隐式RGB帧观察转变为显式运动空间诊断。通过构建运动感知表示,MotionInsight使细微的运动缺陷更加可观察。我们还在GRPO期间引入了运动特定奖励,以将观察到的运动转化为诊断评估。实验表明,MotionInsight为诊断物体运动缺陷提供了有效基础,在三个维度上产生了与人类对齐的得分和基于依据的解释。
英文摘要
Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.