arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sci-VBench:面向科学领域知识与推理密集型视频生成的评估

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao

arXiv 2608.09873首次发表:更新:

发表机构

Zhejiang University; UCAS; Tongji University; Yale University(浙江大学; 中国科学院大学; 同济大学; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出Sci-VBench基准,评估科学领域视频生成,测试16个模型后发现,视觉真实感提升未转化为科学与因果动态建模能力,且专有模型性能优于开源模型。

AI 中文摘要

我们推出Sci-VBench,这是一个用于评估跨科学领域知识与推理密集型视频生成的综合基准。它包含1253个经专家标注的示例,涵盖自然科学、医疗保健、人文与社会科学、工程四大核心学科的60个主题。每个示例要求模型生成具有时间丰富性的视频,这些视频需要科学推理和基于知识的合成,超越表面层面的视觉合理性。我们还建立了基于评分标准的评估方案。我们的分析表明,在该方案下,非专家人类评估者和多模态大语言模型(MLLM)作为评判系统都能与专家判断达成相对较高的一致性,支持大规模可复现评估。我们对16个前沿的专有和开源模型进行了基准测试,发现尽管各系统的自动感知质量得分紧密聚集,但在提示词接地、科学正确性和因果正确性方面的性能差异显著,存在明显的专有-开源差距。这些发现表明,视觉真实感的进步尚未转化为对科学和因果动态的可靠建模。

英文摘要

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.

CommentsCOLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑