AI 中文总结
研究关键帧条件视频生成,提出关键帧指南针综合基准及自动评估框架,将关键帧执行分解为六个指标评估,通过实验揭示当前模型在关键帧执行与自然视频合成间存在权衡等局限性。
AI 中文摘要
视频生成越来越依赖基于关键帧的工作流程,创作者指定参考图像序列来引导生成。尽管近期模型支持多关键帧条件,但它们能否在保持整体视频质量的同时忠实地再现规定关键帧仍不明确。我们提出关键帧指南针,首个评估关键帧条件视频生成的综合基准。它包含386个精心策划样本,涵盖多个方面。还引入自动评估框架,将关键帧执行分解为六个指标,通过专门感知模型增强的基于证据的MLLM判断评估整体视频质量。实验揭示了当前模型的一些基本局限性。
英文摘要
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.
Comments35 pages