arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PragMatch:分离大视觉语言模型中的语用不一致与跨模态不匹配

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh

arXiv 2608.09772首次发表:更新:

发表机构

Saarland University; Max Planck Institute for Informatics(萨尔大学; 马克斯·普朗克信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出PragMatch基准,揭示大视觉语言模型易受表面线索影响,为评估多模态语用推理提供测试平台。

AI 中文摘要

大视觉语言模型(LVLMs)在多模态基准测试中展现出强大性能,但目前仍不清楚它们是真正推理图像与文本间的关系,还是依赖于被称为捷径学习的表面关联。这一问题对多模态讽刺检测尤为重要,因为成功预测依赖于识别语用不一致,而非将讽刺视为简单的图像-文本不匹配。我们推出PragMatch,这是一个源自MMSD2.0的包含3000个图像-文本对的受控基准,涵盖原始讽刺示例及构造的字面与硬负样本对。我们通过系统性掩码识别有影响力的捷径线索,并通过针对性注入实验评估其影响。结果显示,LVLM的预测对词汇、OCR衍生及风格线索敏感,尽管图像-文本关系未变,注入的表面信号仍会导致模型预测发生重大变化。我们的发现揭示了当前LVLMs的局限性,同时PragMatch为评估超越表面图像-文本对齐的多模态语用推理提供了系统性测试平台。

英文摘要

Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations, known as shortcut learning. This question is particularly important for multimodal sarcasm detection, where successful prediction depends on recognizing pragmatic incongruity rather than treating sarcasm as simple image-text mismatch. We introduce PragMatch, a controlled benchmark of 3,000 image-text pairs derived from MMSD2.0, including original sarcastic examples and constructed literal and hard-negative pairs. We identify influential shortcut cues through systematic masking and evaluate their impact through targeted injection experiments. Our results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships. Our findings reveal limitations in current LVLMs while PragMatch provides a systematic testbed for evaluating multimodal pragmatic reasoning beyond surface-level image-text alignment.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑