arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30068cs.CV

TAKE 85:基于85小时电影时长的电影创作者意图测试基准

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

  • LIX, Ecole Polytechnique, IP Paris(巴黎综合理工学院LIX实验室、巴黎IP大学)
  • New York University(纽约大学)
  • Ecole nationale supérieure Louis-Lumière(路易卢米耶高等国家学院)
  • Stony Brook University(石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang, Dimitris Samaras, Vicky Kalogeiton

AI总结:

本文推出首个导演意图理解基准TAKE 85,通过实验发现当前顶尖MLLMs在感知识别与意图理解间存在显著差距,最强模型仅得58分,单一模态无法支撑意图理解。

AI中文摘要:

电影通过灯光、色彩、构图、剪辑、对白、音乐及音效等刻意的创作选择进行传达,人类会自然地将这些信号解读为导演意图,但当前多模态大语言模型(MLLMs)的评估几乎仅聚焦于理解电影中发生了什么,而非为何以这种方式呈现。本文推出TAKE 85,这是首个用于导演意图理解的基准,包含398部短片(总时长85小时),配有专家验证的问答对,涵盖全局及细粒度的视觉与音频意图。通过受控的模态消融实验,TAKE 85可实现对多模态推理的系统性评估。对当前最先进的MLLMs开展的实验显示,感知识别与意图理解之间存在显著差距:尽管模型能准确描述事件与叙事,但始终无法推断出电影创作决策的传达作用。本文的结果确立了导演意图是多模态理解中此前被忽视的维度:即便是最强的模型也仅达到100分中的58分,且消融实验表明,没有任何单一输入模态足以支撑意图理解。所有代码、问答对及模型均可通过该https URL公开获取。

英文摘要:

Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal large language models (MLLMs) are evaluated almost exclusively on understanding what happens rather than why it is presented that way. We introduce TAKE 85, the first benchmark for directorial-intent understanding, comprising 398 short films (85 hours) with expert-verified question-answer pairs spanning global and fine-grained visual and audio intent. Through controlled modality ablations, TAKE 85 enables systematic evaluation of multimodal reasoning. Experiments on state-of-the-art MLLMs reveal a substantial gap between perceptual recognition and intentional understanding: while models accurately describe events and narratives, they consistently fail to infer the communicative role of filmmaking decisions. Our results establish directorial intent as a previously overlooked dimension of multimodal understanding: even the strongest model reaches only 58 out of 100, and our ablations show that no input modality is sufficient on its own. All code, Q&As, and models are publicly available from https://github.com/KaiShinozakiConefrey/Take-85

补充信息

↑