arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27127cs.AI

TransMeme:用于跨文化梗图再创作的多智能体框架

TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation

  • Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Wuhan University(武汉大学)
  • Aarhus University(奥胡斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He

AI总结:

本研究针对跨文化梗图再创作的三大核心挑战,提出多智能体框架TransMeme,经中英双向梗图再创作的人工评估与LLM评判,性能优于所有基线,为该领域未来幽默迁移研究指明方向。

AI中文摘要:

网络梗图是一种普遍存在的多模态在线交流形式,但这类交流常涉及来自不同语言和文化背景的用户。因此,跨文化和跨语言的梗图再创作是实现在线交流中相互理解的核心挑战。与普通翻译或独立的文本改写不同,跨文化梗图再创作必须同时保留交流意图、为目标受众调整文化依赖的含义,并维持文本与图像之间的连贯性。在本研究中,我们首先对跨文化梗图再创作进行明确的任务分析,确定了三个核心挑战:文化特定知识理解、意图与语气保留以及多模态一致性。基于该分析,我们提出了一种多智能体框架,该框架配备了专门的智能体,通过文化适配、目标文本改写、修订和条件视觉调整来协调应对这些挑战。该框架通过协调反馈强化目标文本适配,以处理需要更深入文化或视觉干预的困难案例。我们在中英双向梗图再创作上对该框架进行了评估,采用了人工评估和大语言模型(LLM)作为评判者两种方式。我们的方法在两种评估设置下均始终优于所有基线。在人工评估中,它在所有四个维度上均取得最佳性能,且比最强基线平均提升了33.1%;在LLM作为评判者的评估中,它获得了最高的Top-1排名率(60%,而次优基线为26%)。进一步分析表明,每个组件都对性能有贡献。我们的错误分析显示,剩余瓶颈在于幽默重构和图文对齐,而非简单的文化知识缺口,这为未来关于幽默迁移的研究指明了方向。

英文摘要:

Internet memes are a pervasive form of multimodal online communication; however, such communication often involves users from diverse linguistic and cultural backgrounds. Therefore, adapting memes across cultures and languages is a central challenge for enabling mutual understanding in online communication. Unlike ordinary translation or standalone text rewriting, cross-cultural meme transcreation must jointly preserve communicative intent, adapt culture-dependent meaning for the target audience, and maintain coherence between text and image. In this work, we first provide an explicit task analysis of cross-cultural meme transcreation and identify three core challenges: culture-specific knowledge understanding, intent and tone preservation, and multimodal consistency. Based on this analysis, we propose a multi-agent framework with specialized agents that are coordinated to address these challenges through cultural adaptation, target text rewriting, revision, and conditional visual adjustment. The framework strengthens target text adaptation with coordinated feedback to handle difficult cases that require deeper cultural or visual intervention. We evaluate the framework on bidirectional Chinese-English meme transcreation using both human evaluation and LLM-as-a-Judge. Our method consistently outperforms all baselines across both evaluation settings. In human evaluation, it achieves the best performance on all four dimensions and delivers a 33.1% average improvement over the strongest baseline, while in LLM-as-a-Judge, it attains the highest Top-1 ranking rate (60% versus 26% for the second-best baseline). Further analysis indicates that each component contributes to the performance. Our error analysis suggests that the remaining bottlenecks lie in humor reconstruction and image-text alignment rather than simple cultural knowledge gaps, pointing to future work on humor transfer.

补充信息

↑