发表机构
Arizona State University; Capital One(亚利桑那州立大学; 第一资本金融公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建了多模态学术对话基准SciReC,提出缺陷诊断框架DMRA评估多模态大语言模型的关系推理能力,发现不同模型在各类关系及领域上的表现差异,明确错误的主要来源。
AI 中文摘要
关系推理需要完成感知理解、比较以及整合概念间潜在关系的过程,该能力包含类比、结构、因果等多个类别,分别对应高阶理解的不同维度。为检验多模态大语言模型(MLLM)在这些关系推理任务上的表现,我们构建了SciReC,这是一个模型自适应的多模态学术对话基准。由于关系推理过程涉及多种表征及视觉理解、知识呈现、记忆回忆等不同因素,我们提出了DMRA,这是一种基于缺陷的诊断框架,可量化各组件的贡献以识别失败案例的核心原因。Claude 4.6在整体关系得分上取得73%的最佳表现,GPT 5.4以68%紧随其后。性能趋势显示,开源模型在空间关系上得分最低,而专有模型在层级和序列关系上表现更差;跨领域来看,模型在天文学领域得分最低,在心理学领域得分最高。DMRA的结果表明,关系推理是所有模型的主要错误来源,其次是记忆限制。
英文摘要
Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding. To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal academic dialog benchmark. As the relational reasoning process involves multiple representations and various factors (visual understanding, exhibiting knowledge, and memory recall), we propose DMRA, a deficit-based diagnostic framework that quantifies the contribution of these components to identify the primary cause of unsuccessful cases. Claude 4.6 achieved the best performance on the overall relational score with 73\%, followed by GPT 5.4 with 68\%. Performance trends indicate that open-source models achieve their lowest scores on spatial relations, while proprietary models struggle more with hierarchical and sequential relations. Across domains, model performance is lowest on Astronomy and highest on Psychology. The results of DMRA reveal that relational reasoning is the primary source of error across all models, followed by memory limitations.