发表机构
Texas A&M University; Purdue University; University of Wisconsin–Madison(德克萨斯农工大学; 普渡大学; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了针对美国起源梗图的双向文化数据集MemeBridge,用于测试跨文化理解,发现LLM跨文化理解能力不足,用MemeBridge微调可提升模型性能。
AI 中文摘要
跨文化交流本身就具有挑战性,尤其是通过梗图这类文化密集且模糊的形式进行交流。尽管人们期望大型语言模型(LLM)有望弥合此类文化差异,但现有的基准数据集往往无法捕捉准确理解所需的文化背景。为解决这一问题,我们引入MemeBridge,这一以美国起源梗图为核心的精选数据集,旨在捕捉两种互补视角:(1)中国参与者对这些梗图的理解;(2)美国参与者对其他文化背景人群可能产生的误解的预判。此处的背景指隐性文化知识,包括塑造梗图理解的背景信念、规范和共同假设。该数据集通过多阶段众包流程构建,并经过严格验证,包括人工一致性检查和基于GPT的分类验证。每个梗图都被标注了情感、情绪、文化意义和知识类型,为下游任务提供丰富的监督信号。值得注意的是,我们发现美国参与者预判的误解往往不准确,凸显了文化理解中的不对称性以及采用自身之外视角的挑战。这种聚焦表达与感知的双向框架,能够对跨文化理解进行更细致的基准测试。我们对多个LLM的探测显示,尽管不同文化背景下开发的模型展现出部分跨文化理解能力,但往往难以进行复杂的理解。相比之下,使用MemeBridge进行微调可提升模型性能,凸显了基于文化的资源在全球多样化环境中训练和评估LLM的价值。
英文摘要
Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect large language models (LLMs) to hold promise for bridging such gaps, existing benchmark datasets often fail to capture the cultural context necessary for accurate interpretation. To address this, we introduce MemeBridge, a curated dataset centered on U.S.-originated memes, designed to capture two complementary perspectives: (1) how Chinese participants interpret these memes, and (2) how U.S. participants anticipate how people from other cultures might misunderstand them. Here, context refers to implicit cultural knowledge, including background beliefs, norms, and shared assumptions that shape meme comprehension. The dataset was constructed via a multi-stage crowdsourcing pipeline with rigorous validation, including human agreement checks and GPT-based classification verification. Each meme is annotated with sentiment, emotion, cultural significance, and knowledge type, providing rich supervision for downstream tasks. Notably, we observe that the anticipated misunderstandings from U.S. participants are often inaccurate, highlighting the asymmetries in cultural understanding and the challenges of adopting perspectives beyond one's own. This bidirectional framing, which focuses on both expression and perception, enables more nuanced benchmarking of cross-cultural comprehension. Our probing of multiple LLMs reveals that while models developed in different cultural contexts exhibit partial cross-cultural understanding, they often struggle with sophisticated interpretations. By contrast, fine-tuning with MemeBridge improves model performance, underscoring the value of culturally grounded resources for training and evaluating LLMs in globally diverse settings.
Journal refIn Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Vol. 1 (KDD '26), 2026