发表机构
Renmin University of China; Alibaba Group; Tianjin University(中国人民大学; 阿里巴巴集团; 天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DiVid提出维度级诊断框架,分解视频生成多样性为六个维度,揭示模型在运动和相机上的特定崩溃,并识别默认模式收敛与实现差距两大瓶颈,为维度感知训练提供方向。
AI 中文摘要
尽管取得了显著进展,视频生成模型在从同一提示词重复采样时,往往会产生高度相似的输出,这限制了其在创意探索中的实用性。现有的多样性评估主要依赖全局标量指标,这掩盖了视频时空空间中多样性崩溃的具体位置。我们提出了DiVid,一个维度级诊断框架,将视频生成多样性分解为六个可解释的维度:语义、风格、主体、场景、运动和相机。每个维度通过可复现的计算机视觉流程进行测量,并与质量和指令忠实度分析相结合,以考察潜在的权衡。对代表性视频生成模型的系统评估显示,多样性具有高度维度特异性:具有强全局多样性分数的模型仍然会在特定因素上崩溃,尤其是运动和相机。这些排名在过滤不忠实生成后依然存在,表明这是真实的能力差异,而非偏离提示的输出。在测量之外,受控提示干预识别出两个根本瓶颈:默认模式收敛,即模型在开放式提示下退回到主导模式;以及实现差距,即模型未能忠实实现多样化的、明确请求的替代方案,尤其是对于时间因素。时间因素的更大忠实度损失凸显了仅通过文本控制运动和相机变化的难度。因此,DiVid将多样性研究从测量其是否存在,转向诊断其在何处及为何崩溃,并为维度感知的训练目标和控制信号提供了可行的方向。该框架将发布,以促进未来关于多样化和可控视频生成的研究。
英文摘要
Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rely on global scalar metrics, which obscure where diversity collapses in the spatiotemporal space of videos. We introduce DiVid, a dimension-level diagnostic framework that decomposes video generation diversity into six interpretable dimensions: Semantic, Style, Subject, Scene, Motion, and Camera. Each dimension is measured through a reproducible computer-vision pipeline and analyzed alongside quality and instruction faithfulness to examine potential trade-offs. Systematic evaluation of representative video generation models reveals that diversity is highly dimension-specific: models with strong global diversity scores still collapse on specific factors, particularly Motion and Camera. These rankings persist after filtering unfaithful generations, indicating genuine capability differences rather than off-prompt outputs. Beyond measurement, controlled prompt interventions identify two fundamental bottlenecks: default mode convergence, where models fall back to dominant patterns under open-ended prompts; and realization gaps, where models fail to faithfully realize diverse, explicitly requested alternatives, particularly for temporal factors. The larger faithfulness losses for temporal factors highlight the difficulty of controlling motion and camera variation through text alone. DiVid thus shifts the study of diversity from measuring whether it exists to diagnosing where and why it collapses, and provides actionable directions for dimension-aware training objectives and control signals. The framework will be released to facilitate future research on diverse and controllable video generation.