arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05817cs.CL

M$^3$R-Bench:一个用于基于证据的多模态隐喻理解的统一基准

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

Hong Jiang, Junnan Zhu, Jingwang Huang, Xiao Sun, Yuming Yang, Jiang Zhong, Ruirui Chen, Jingman Shi, Hao Wu, Nayu Liu, Xinyi Jiang, Kaiwen Wei

中文总结 AI 辅助

本文提出多模态隐喻理解基准M$^3$R-Bench,针对现有模型的跨模态证据-映射不匹配问题,构建M$^3$R-Reasoner方法,在多指标上超越GPT-5.5等现有模型。

中文摘要 AI 辅助

隐喻通过跨域映射实现对抽象概念的理解,同时传递情感态度。在多模态场景中,视觉与文本信息共同构建目标-源映射,要求模型同时具备概念理解与跨模态推理能力。然而,现有基准主要通过孤立子任务评估隐喻理解,且缺乏基于证据的解释,难以评估模型是否建立了基于视觉与文本的映射。为解决这些局限,本文引入M$^3$R-Bench,这是一个统一且基于证据的基准,包含1000个人工验证标注的图像-文本实例。在概念隐喻理论与非字面语言理解理论的指导下,M$^3$R-Bench提供隐喻存在、目标-源映射、情感以及遵循“证据识别-映射建立-情感推理”的阶段性解释的联合标注。对M$^3$R-Bench的评估显示,现有模型常忽略视觉证据、依赖表面文本线索并生成不准确的目标-源映射,暴露出跨模态证据-映射不匹配问题。为解决该不匹配,本文提出M$^3$R-Reasoner,其结合基于课程的推理监督与任务感知强化学习,使模型推理与隐喻解释对齐。实验表明,仅使用8B参数骨干的M$^3$R-Reasoner在四项统一任务指标上优于更大的专有多模态大语言模型(MLLMs),视觉证据与情感理由得分分别较GPT-5.5提升28.45与30.11个百分点,在平均 rubric 得分上超越Claude-Sonnet-4.6达8.00个百分点。该数据集与代码可在指定网址获取。

英文摘要

Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues.To address these limitations, we introduce M$^3$R-Bench, a unified and evidence-grounded benchmark containing 1,000 image--text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M$^3$R-Bench provides joint annotations for metaphor occurrence, Target--Source mapping, sentiment, and stage-wise explanations following ``evidence identification--mapping establishment--sentiment inference.''Evaluations on M$^3$R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target--Source mappings, exposing a cross-modal evidence--mapping mismatch. To address this mismatch, we propose M$^3$R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M$^3$R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at https://github.com/hongshi4/M3R-Bench.

补充信息

↑