arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CinematicVQA:大型视觉语言模型中的电影语法推理基准

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

Shuo Xing, Pooja Verlani, Balu Adsumilli, Zhengzhong Tu

arXiv 2609.28813首次发表:更新:

发表机构

Texas A&M University; Google Inc.(德克萨斯A&M大学; 谷歌公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CinematicVQA基准,利用电影场景图评估LVLMs的电影语法推理,揭示模型在描述视觉呈现优于识别技术,微调可提升叙事功能与多跳推理。

AI 中文摘要

电影摄影,即通过取景、灯光和摄像机操作进行视觉叙事的技艺,从根本上塑造了观众感知视频内容并与之产生情感共鸣的方式。尽管大型视觉语言模型(LVLMs)在视频问答方面取得了显著进展,但现有基准主要侧重于识别低级技术,而非理解其叙事影响。为解决这一问题,我们引入了CinematicVQA,这是首个超越技术识别、评估电影语法推理的电影视频理解基准,利用我们提出的电影场景图(CSG),一种将拍摄技术与其感知效果和叙事功能联系起来的结构化表示。通过对最先进的LVLMs进行全面评估,我们揭示了一个显著的语义鸿沟:模型在描述视觉呈现方面的表现始终高于识别底层技术。令人惊讶的是,思维链提示未能提供一致的改进,反而降低了大多数模型的性能,这表明当前的LVLMs缺乏足够的电影领域知识,无法从逐步推理中受益。在CinematicVQA-train上进行微调带来了持续的改进,尤其是在叙事功能和多跳推理方面。总体而言,CinematicVQA既作为LVLMs中电影评估的严格基准,也作为训练更具电影感知能力的视频模型的实用数据集。

英文摘要

Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.

Comments6 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑