发表机构
Singapore Management University; VNU-HCM(新加坡管理大学; 胡志明市国家大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出东南亚文化时刻基准CMB,通过三阶段评估视觉-语言模型的文化理解能力,发现模型跨阶段能力未完全级联,音频在非拉丁文字国家多为干扰,专家对邻国概念的识别也不及随机水平。
AI 中文摘要
视频中的文化理解不仅是识别可见内容,还需掌握文化概念的象征意义与时间意义。我们将其分解为三种能力:命名概念的象征意义、在视频中视觉识别该概念、在时间上定位其子事件。现有视频文化基准倾向于测试可见内容,将这三种能力合并为单一分数,掩盖了瓶颈。我们推出文化时刻基准(Cultural Moment Benchmark, CMB):来自东南亚7个国家、5个类别的306个由专家筛选的概念。我们通过三个阶段评估每个概念,每个阶段对应一种能力:给定描述,阶段1(S1)从4个候选概念名称中选择,阶段2(S2)从4个候选视频时刻中选择,阶段3(S3)预测该时刻在视频中的起止时间。为保持每个阶段专注于不同能力,我们采用三项设计选择:语义相似干扰项(S1、S2)、未标记视频时刻(S2)、在不同示例视频上的自由形式定位(S3)。在6个视觉-语言模型中,失败模式因能力和模态而异:i)即便是最强的闭源模型,当三个阶段都需正确时,得分仍低于30%;ii)三种能力并非完全级联:正确命名概念对一半模型的视频识别有帮助,但视频识别对时间定位几乎无影响;iii)音频根据概念的不同,起互补、冗余或干扰作用,在非拉丁文字国家更常为干扰项,同时移除音频和字幕对游戏与音乐的伤害最大。我们的14名评分者的人类研究显示,即便是专家评分者对邻国概念的得分也低于随机水平,表明CMB需要特定国家的文化知识。CMB作为诊断工具,将失败归因于特定能力或模态。
英文摘要
Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recognizing it on video, and locating its sub-events in time. Existing video-cultural benchmarks tend to test what is seen, collapsing these three abilities into a single score that hides the bottleneck. We introduce the Cultural Moment Benchmark (CMB): 306 expert-curated concepts from seven countries in Southeast Asia across five categories. We evaluate each concept through three stages, one per ability. Given a description, Stage 1 (S1) selects from four candidate concept names, Stage 2 (S2) selects from four candidate video moments, and Stage 3 (S3) predicts the start and end times of the moment in a video. To keep each stage focused on a distinct ability, we use three design choices: semantic-similarity distractors (S1, S2), unlabeled video moments (S2), and free-form localization on a different example video (S3). Across six vision-language models, failure modes vary by ability and modality. i) Even the strongest closed-source models score below 30% when all three stages must be correct; ii) The three abilities do not fully cascade: naming a concept correctly helps half the models recognize it on video, but recognizing it has little effect on locating the sub-event in time; iii) Audio is complementary, redundant, or distracting depending on the concept, more often distracting in non-Latin-script countries; removing both audio and subtitles hurts Games and Music the most. Our 14-rater human study shows that even Expert raters score below chance on concepts from a neighboring country, indicating that CMB requires country-specific cultural knowledge. CMB acts as a diagnostic harness, attributing failures to a specific ability or modality.
CommentsAccepted to EMNLP 2026 Main Conference, https://culturalmoment-benchmark.github.io/