arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemeBench:大型视觉语言模型在解读依赖文化的梗图时所缺失的内容

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

Weihang Wang, Kainan Tu, Jielei Zhang, Run Yang, Boheng Sheng, Yuchen He, Yu Xie, Pengyu Chen, Peiyi Li, Huyang Sun, Longwen Gao, Zhouhui Lian

arXiv 2607.27798首次发表:更新:

AI 中文总结

研究针对LVLMs解读文化类梗图的知识缺口问题,推出诊断基准MemeBench及实体引导检索基线KAR,验证了检索手段可提升梗图解读的VIKR成功率。

AI 中文摘要

大型视觉语言模型(LVLMs)在描述视觉内容方面已有进步,但准确的描述并不能保证当含义依赖于像素之外的知识时能完成解读。梗图(memes)凸显了这一差距,因为它们依赖文化实体、背景知识和社区惯例。大多数梗图基准将解读简化为标签或整体分数,掩盖了解释失败的具体环节。我们推出MemeBench,这是一个包含1253张中英文梗图的诊断基准,配有人工编写的参考资料和质量受控的VIKR标注,核心围绕动漫、漫画、游戏及相关网络亚文化。其VIKR模式将解释分解为视觉线索(Visual clues)、身份关联(Identity links)、知识单元(Knowledge units)和推理机制(Reasoning mechanisms)。在26个LVLMs中,所有模型对可见内容的覆盖都比对解读所需知识的覆盖更可靠,即使是最强的模型也存在22.6%的视觉-知识差距。为验证该诊断能否指导改进,我们推出KAR,这是一个基于CultureBase构建的实体引导检索基线。在四个受控模型中,KAR将VIKR成功率提升了3.6%至7.4%,与通用检索相比,它修复了更多答案且中断更少。不过,两种检索条件在所有比较中都提升了身份和知识相关表现,同时降低了视觉覆盖度。MemeBench可判断解读是否成功、缺失了什么,以及针对性证据能否填补已诊断的差距。

英文摘要

Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.

Comments17 pages, 5 figures, and 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑