评估热门好莱坞电影的多模态叙事理解能力
Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
- Bowdoin College(鲍登学院)
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究构建了符合票房受欢迎度与公有领域标准的好莱坞电影多模态数据集,建立了相关MCQ基准,发现多数视觉-语言模型表现接近随机,视听模型最高准确率61.1%,均低于人类水平。
AI中文摘要:
多模态语言模型在实现电影的大规模计算分析方面展现出越来越大的潜力,为了解电影史和叙事技巧的演变开辟了新途径。但围绕好莱坞电影构建稳定基准的工作因版权保护而变得复杂。在本研究中,我们直接解决这些问题,依据两个标准构建了一套新的好莱坞电影集合:票房受欢迎度(我们发布了1922年至1979年《综艺》杂志报道的首个大规模公开周度票房收入集合);以及可能的公有领域状态(通过查询美国版权登记目录中的版权登记和续期信息确定)。我们在该集合的基础上构建了一个新的多模态多项选择(MCQ)基准,该基准聚焦于叙事元素,直接评估模型为电影叙事的有意义研究提供信息的能力;我们发现许多视觉-语言模型在该任务上表现不佳(许多模型的准确率接近随机水平),而视听模型(包括那些在场景字幕中使用音频的模型)的最高准确率为61.1%,远低于人类水平的表现。
英文摘要:
Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-language models struggle on this task (with many performing at near-chance levels of accuracy), while audio-visual models (including those that use audio in captioning scenes) reach a maximum accuracy of 61.1%, well below human-level performance.