arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36218cs.CLcs.AIcs.ETcs.LG

CineSubBench:从多语言电影字幕评估大语言模型的长篇叙事与文化理解能力

CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles

Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei

首次发表
浏览论文内容

中文总结 AI 辅助

提出CineSubBench基准,利用多语言电影字幕评估LLM的长篇叙事理解、多语言一致性及文化分级能力,涵盖1012部电影、7项任务,发现模型在情节前提恢复上优于完整概要,且跨语言与文化校准存在显著差异。

中文摘要 AI 辅助

大语言模型在法律、医学、软件工程和网络安全等专业领域的评估日益增多,然而电影领域相对而言仍未得到充分探索,尽管电影需要长篇叙事整合、多语言解读以及基于文化背景的受众判断。我们推出了CineSubBench,一个用于从多语言电影字幕评估长上下文电影理解的基准测试。一条字幕轨道将一部电影表示为数千条按时间顺序排列的简短话语,模型必须从中重建角色、关系、事件、因果推进和主题,而无需明确的场景或事件结构。CineSubBench包含1,012部电影,涵盖六种语言的完整字幕,共生成6,072条字幕轨道和813万条带时间戳的字幕条目。它提供了一个匹配的多任务、多语言和多文化(MultiX)评估环境:七项任务涵盖叙事重建与抽象、类型预测、年龄适宜性、跨越十个国家分级系统的特定国家电影分级,以及基于字幕的语言安全性。在九个LLM中,情节前提的恢复比事件完整的概要更可靠;跨语言一致性在不同模型和语言之间差异显著;国家分级系统揭示了不同的校准模式;强亵渎语言比轻度粗俗语言更容易被证实。CineSubBench将电影确立为长上下文LLM评估领域,并提供了一个用于衡量叙事、多语言、文化和证据基础能力的统一基准。

英文摘要

Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6,072 tracks and 8.13M timestamped subtitle entries. It provides a matched multi-task, multilingual, and multicultural (MultiX) evaluation setting: seven tasks span narrative reconstruction and abstraction, genre prediction, age suitability, country-specific motion-picture ratings across ten national classification systems, and subtitle-grounded language safety. Across nine LLMs, plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies substantially across models and languages; national rating systems expose distinct calibration patterns; and strong profanity is far easier to ground than mild obscenity. CineSubBench establishes film as a long-context LLM evaluation domain and provides a unified benchmark for measuring narrative, multilingual, cultural, and evidence-grounding capabilities.

发表机构

  • University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑