基于多模态语言模型的计算幽默:方法、数据集、评估及挑战
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
- Case Western Reserve University(凯斯西储大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究聚焦于单图像和多面板作品中的视觉幽默理解,通过与先前综述对比,按能力层次组织文献,综合基准设计等内容,追溯领域转变,指出进展的主要障碍包括易出现捷径的评估、文化覆盖有限等问题。
AI中文摘要:
对于人工智能系统来说,理解表情包、漫画和连环漫画中的多模态幽默仍然很困难,因为其含义依赖于非字面机制、共享文化知识和交际意图,而非字面场景描述。本综述聚焦于单图像和多面板作品中的视觉幽默理解,同时将幽默生成视为新兴的下游前沿领域。我们将文献与先前的幽默、讽刺和通用多模态语言模型综述进行对比,并按以能力为中心的层次结构进行组织,涵盖识别、解释与推理以及生成。在此视角下,我们综合了基准设计、评估协议和建模范式,追溯了该领域从特定任务融合模型到基于多模态对齐、基于证据的推理和可控生成的大模型方法的转变。我们通过强调进展的主要障碍得出结论:易于出现捷径的评估、有限的文化和叙事覆盖范围、薄弱的证据基础以及未解决的安全和所有权问题。
英文摘要:
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.