发表机构
Microsoft Research India; LinkedIn; Nutanix; Indian Institute of Science; Apple; Fujitsu Research India; Indian Institute of Technology Kharagpur(微软研究院印度分部; 领英公司; Nutanix公司; 印度科学学院; 苹果公司; 富士通印度研究院; 印度克勒格布尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CaRGo-T因果推理思维图框架,在四类含讽刺等喜剧内容的数据集上,提升VLMs的幽默理解与检测性能,相关代码已公开。
AI 中文摘要
大规模视觉语言模型(VLMs)已在各类多模态任务中展现出出色的通用性,但理解幽默内容仍具挑战性,因为幽默内容往往依赖图像与文本模态间实体、事件、上下文及隐含关系的微妙交互,这类交互涉及复杂推理链,难以通过传统提示或线性思维链推理捕捉。本研究提出CaRGo-T(因果推理思维图,Causal Reasoning Graph-of-Thought),该推理框架将多模态幽默背后的因果与上下文关系表示为轻量级图式推理结构。此图被序列化为由VLM生成的基于代码的表示,后续可由相同或不同VLM解释,以零样本或上下文学习设置下生成最终预测。我们在涵盖讽刺、挖苦、梗图等不同喜剧内容形式的四个数据集上,对CaRGo-T进行幽默理解与幽默检测评估。对最先进的商用及开源VLMs的实验显示,CaRGo-T相较现有基于推理的基线方法,性能持续提升,在幽默理解任务上实现约1%-20%的增益,在幽默检测任务上实现约1%-3%的增益。进一步的互信息分析表明,CaRGo-T生成的推理表示包含比基线推理方法更多与目标输出相关的信息。代码可访问此https URL获取。
英文摘要
Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humor remains challenging because humorous content often depends on subtle interactions among entities, events, context, and implicit relationships across image and text modalities. These interactions can involve complex chains of reasoning that are difficult to capture through conventional prompting or linear chain-of-thought reasoning. In this work, we propose CaRGo-T (Causal Reasoning Graph-of-Thought), a reasoning framework that represents the causal and contextual relationships underlying multimodal humor as a lightweight graph-based reasoning structure. The graph is serialized into a code-based representation generated by a VLM, which can subsequently be interpreted by the same or a different VLM to produce the final prediction in zero-shot or in-context learning settings. We evaluate CaRGo-T on humor understanding and humor detection across four datasets spanning diverse forms of comedic content, including satire, sarcasm, and memes. Experiments with state-of-the-art commercial and open-source VLMs show that CaRGo-T consistently improves performance over existing reasoning-based baselines, achieving gains of approximately 1-20% on humor understanding and 1-3% on humor detection. Further analysis using mutual information indicates that the reasoning representations produced by CaRGo-T contain more information relevant to the target output than those generated by baseline reasoning approaches. Code is available at https://github.com/abhi1nandy2/CaRGo-T.
Comments18 pages, 5 figures