arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CultureVidBench:文本到视频生成中的文化理解基准测试

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu

arXiv 2608.01942首次发表:更新:

发表机构

Nanyang Technological University; Centrale Supélec; Singapore Management University; University of Cambridge(南洋理工大学; 中央高等电力学院; 新加坡管理大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出CultureVidBench基准,评估7款T2V模型的文化理解能力,发现现有模型难捕捉细粒度文化细节,尤其针对代表性不足的文化区域与线索。

AI 中文摘要

文本到视频(T2V)生成模型发展迅速,但它们对多元文化语境的表征能力仍未得到充分探索。现有基准主要关注感知质量、物理合理性和文本-视频对齐,却未直接评估生成视频是否捕捉到文化特有的物体、动作、仪式、可见文本或音频线索。我们推出CultureVidBench,这是一个用于评估T2V生成中文化理解的综合基准。CultureVidBench包含1000个精心策划的提示,覆盖12个国家、6个大洲、8个文化区域以及14个文化方面,分为物质文化、社会实践与表演、仪式与典礼三类。CultureVidBench专为视频生成设计,强调动态和多模态文化表征,包括社会互动、仪式流程以及符合文化规范的可见文本和音频。我们通过用户人工研究和基于多模态大语言模型(MLLM)的自动评估,从文化忠实度、多模态文化呈现、语义贴合度和感知质量四个维度评估了7个具有代表性的T2V模型。结果显示,尽管当前模型在语义贴合度和视觉质量上表现出色,但往往无法忠实地捕捉细粒度的文化细节,尤其是在代表性不足的地区、仪式以及多模态文化线索方面。

英文摘要

Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.

CommentsEMNLP 2026 Main Conference, Project page:https://hanxjing.github.io/CultureVidBench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑