M$^3$R-Bench:一个用于基于证据的多模态隐喻理解的统一基准
M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
中文总结 AI 辅助
本文提出多模态隐喻理解基准M$^3$R-Bench,针对现有模型的跨模态证据-映射不匹配问题,构建M$^3$R-Reasoner方法,在多指标上超越GPT-5.5等现有模型。
中文摘要 AI 辅助
隐喻通过跨域映射实现对抽象概念的理解,同时传递情感态度。在多模态场景中,视觉与文本信息共同构建目标-源映射,要求模型同时具备概念理解与跨模态推理能力。然而,现有基准主要通过孤立子任务评估隐喻理解,且缺乏基于证据的解释,难以评估模型是否建立了基于视觉与文本的映射。为解决这些局限,本文引入M$^3$R-Bench,这是一个统一且基于证据的基准,包含1000个人工验证标注的图像-文本实例。在概念隐喻理论与非字面语言理解理论的指导下,M$^3$R-Bench提供隐喻存在、目标-源映射、情感以及遵循“证据识别-映射建立-情感推理”的阶段性解释的联合标注。对M$^3$R-Bench的评估显示,现有模型常忽略视觉证据、依赖表面文本线索并生成不准确的目标-源映射,暴露出跨模态证据-映射不匹配问题。为解决该不匹配,本文提出M$^3$R-Reasoner,其结合基于课程的推理监督与任务感知强化学习,使模型推理与隐喻解释对齐。实验表明,仅使用8B参数骨干的M$^3$R-Reasoner在四项统一任务指标上优于更大的专有多模态大语言模型(MLLMs),视觉证据与情感理由得分分别较GPT-5.5提升28.45与30.11个百分点,在平均 rubric 得分上超越Claude-Sonnet-4.6达8.00个百分点。该数据集与代码可在指定网址获取。
英文摘要
Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues.To address these limitations, we introduce M$^3$R-Bench, a unified and evidence-grounded benchmark containing 1,000 image--text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M$^3$R-Bench provides joint annotations for metaphor occurrence, Target--Source mapping, sentiment, and stage-wise explanations following ``evidence identification--mapping establishment--sentiment inference.''Evaluations on M$^3$R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target--Source mappings, exposing a cross-modal evidence--mapping mismatch. To address this mismatch, we propose M$^3$R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M$^3$R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at https://github.com/hongshi4/M3R-Bench.