arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemeCULT-1K:多模态模型南亚文化语境与幽默理解基准

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir, Mueeze Al Mushabbir, Mohammed Saidul Islam, Mir Rayat Imtiaz Hossain, Md Tahmid Rahman Laskar, Sabbir Ahmed

arXiv 2609.01772首次发表:更新:

发表机构

Islamic University of Technology; South East University; Vector Institute; University of British Columbia; York University(伊斯兰科技大学; 东南大学; 矢量研究所; 不列颠哥伦比亚大学; 约克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出包含南亚多语言梗图的MemeCULT-1K基准,评估13种视觉-语言模型的梗图理解能力,发现补充文化语境可提升模型表现,同时揭示不同模型的错误类型差异,凸显文化知识整合的必要性。

AI 中文摘要

梗图(Meme)理解不仅需要识别视觉内容或字面文本,还需要大多数视觉-语言模型仍缺乏的隐含文化知识和语用推理。我们推出MemeCULT-1K,这是一个包含1000张南亚梗图的多语言基准,涵盖孟加拉语、英语和印地语,每张梗图都配有文化语境说明和3个人工撰写的解释,另有54张孟加拉语地区方言梗图作为补充。我们在两种设置下评估了13种流行的视觉-语言模型(VLMs):仅梗图设置和语境感知设置。提供最少的文化语境在所有模型和语言中均带来了持续的提升:平均SBERT相似度从44.6提升至56.4(+11.8),BLEURT从37.3提升至42.3(+5.0),大语言模型(LLM)作为评判者的评分从5分制的2.57提升至3.43(+0.86)。细粒度错误分析显示,闭源模型的主要问题是实体和指代识别错误,而开源模型的瓶颈在于更广泛的文化知识缺口,其中语言和语音层面的错误在两类模型中都是最难以通过语境改善的。这些结果凸显了基于文化的梗图理解的难度,并推动未来开展将显式文化知识整合的研究。我们的数据集和代码可在TawsifDipto17/MemeCULT-1K获取。

英文摘要

Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes. We evaluate thirteen popular Vision Language Models (VLMs) under two settings: meme-only and context-aware. Providing minimal cultural context yields consistent gains across all models and languages: mean SBERT similarity improves from 44.6 to 56.4 (+11.8), BLEURT from 37.3 to 42.3 (+5.0), and LLM-as-a-Judge scores from 2.57 to 3.43 out of 5 (+0.86). Fine-grained error analysis reveals that closed-source models fail mainly on entity and reference misidentification, while open-source models are bottlenecked by broader cultural knowledge gaps, with linguistic and phonological failures proving the most context-resistant across both. These results highlight the difficulty of culturally grounded meme understanding and motivate future work on explicit cultural knowledge integration. Our dataset and code are publicly available at TawsifDipto17/MemeCULT-1K.

CommentsAccepted by EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑