PersonaShot:多镜头视频生成中以人为中心的叙事连续性基准
PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation
浏览论文内容
中文总结 AI 辅助
该研究针对多镜头视频生成中角色叙事连续性评估的不足,推出PersonaShot基准,设计三类特定评估器,发现先进模型感知质量与跨镜头连续性存在差距,评估器与人类专家判断一致性强。
中文摘要 AI 辅助
视频生成正从单镜头片段快速发展至多镜头叙事,其中人类角色是核心叙事锚点。然而,现有基准主要评估角色外观或单镜头质量,未衡量不同镜头间身体与情感状态是否保持连贯,也很少提供针对特定标准的评估方法,尽管身体连续性、面部动态和电影关系需要不同的视觉、时间和关系证据。为解决这些局限,我们推出PersonaShot,首个针对多镜头视频生成中叙事连续性的以人为中心的基准。PersonaShot包含约1000个多镜头片段,以及涵盖身体连续性、情感动态和电影语法的16个指标。1)叙事连续性基准:我们在三个时间层面评估角色连贯性:镜头内状态、跨镜头转换和序列级轨迹。2)人类对齐的专家评估器:我们将大型多模态教师的推理提炼为轻量级的特定标准评估器,每个评估器基于其指标所需的视觉、时间或关系证据,并使其与人类专家判断对齐。3)系统评估与见解:我们的评估揭示了最先进模型的不同能力特征,以及感知质量与跨镜头叙事连续性之间的明显差距。即使视觉上引人注目的视频,在镜头间也常出现身体状态重置、情感突然转变和电影关系断裂。人类研究进一步表明,我们的评估器与专家判断之间具有强一致性。
英文摘要
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality, without measuring whether physical and emotional states remain coherent across cuts. They also rarely provide criterion-specific evaluation methods, although physical continuity, facial dynamics, and cinematic relations require different visual, temporal, and relational evidence. To address these limitations, we introduce PersonaShot, the first person-centric benchmark for narrative continuity in multi-shot video generation. PersonaShot contains approximately 1,000 multi-shot segments and 16 metrics spanning physical continuity, affective dynamics, and cinematic grammar. \textbf{\textit{1)} Narrative Continuity Benchmark:} We evaluate character coherence across three temporal levels: within-shot states, cross-shot transitions, and sequence-level trajectories. \textbf{\textit{2)} Human-Aligned Specialist Evaluators:} We distill reasoning from a large multimodal teacher into lightweight criterion-specific evaluators, each grounded in the visual, temporal, or relational evidence required by its metric, and align them with expert human judgments. \textbf{\textit{3)} Systematic Evaluation and Insights:} Our evaluation reveals distinct capability profiles across state-of-the-art models and a clear gap between perceptual quality and cross-shot narrative continuity. Even visually compelling videos frequently exhibit physical-state resets, abrupt affective shifts, and broken cinematic relations across shots. Human studies further demonstrate strong agreement between our evaluators and expert judgments.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Tencent Youtu Lab(腾讯优图实验室)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。