BanglaMemeX:通过多模态可解释数据集推进孟加拉语文化隐喻图像理解
BanglaMemeX: Advancing Cultural Metaphoric Image Interpretation in Bangla with a Multimodal Explainable Dataset
浏览论文内容
中文总结 AI 辅助
本文提出BanglaMemeX,一个含3000个孟加拉语模因的文化多模态基准,标注多维标签和人工解释,评估现代视觉语言模型,发现其难以解读隐性文化线索,强调需文化感知系统。
中文摘要 AI 辅助
视觉语言模型在多模态基准测试中已取得强劲性能,但其对植根于文化且富含隐喻内容的推理能力仍未得到充分研究。互联网模因呈现了一个具有挑战性的场景,其中意义源于图像、叠加文本、讽刺以及共享的社会文化知识之间的隐性交互,而非字面的视觉识别。这一挑战在孟加拉语等低资源语言中被放大,因为语码混合、风格化脚本和文化特定象征引入了显著的分布偏移。在本工作中,我们引入了BanglaMemeX,一个植根于文化的多模态基准,包含3,000个孟加拉语模因,标注了多维标签(幽默、讽刺、冒犯性、动机意图和整体情感)以及明确描述文本和视觉隐喻的人工撰写的解释。我们系统评估了现代视觉语言模型在分类和解释生成上的表现,揭示当前模型尽管在表面层面具有合理准确性,但仍难以解读隐性文化线索。我们的结果强调了在语言和文化分布偏移下,具备扎根推理能力的文化感知多模态系统的必要性。
英文摘要
Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason about culturally grounded and metaphor-rich content remains insufficiently studied. Internet memes present a challenging setting where meaning emerges from implicit interactions between image, overlaid text, sarcasm, and shared socio-cultural knowledge rather than literal visual recognition. This challenge is amplified in low-resource languages such as Bangla, where code-mixing, stylized scripts, and culturally specific symbolism introduce substantial distribution shift. In this work, we introduce BanglaMemeX, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels (humor, sarcasm, offensiveness, motivational intent, and overall sentiment) and human-written explanations that explicitly describe textual and visual metaphors. We systematically evaluate modern VLMs on both classification and explanation generation, revealing that current models struggle to interpret implicit cultural cues despite reasonable surface-level accuracy. Our results highlight the need for culturally-aware multimodal systems capable of grounded reasoning under linguistic and cultural distribution shift.
发表机构
- University of Dhaka(达卡大学)
- University of Texas at Dallas(德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。