发表机构
IIIT Delhi; BITS Pilani; Microsoft Research India; IIT Kanpur(德里印度理工学院; 皮拉尼比尔拉理工学院; 微软研究院印度分院; 坎普尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出五任务诊断实验,分离多模态LLM在物理与几何推理中的感知与推理失败,发现感知错误导致性能下降,且失败类型因领域而异,并指出InternS1-mini表现不佳。
AI 中文摘要
多模态大语言模型在科学推理基准上报告了强劲的性能,但大多数模型将感知和推理视为单一的可测量过程。我们引入了一个涵盖物理和几何基准的五任务诊断实验,该实验将失败隔离为感知、推理或两者兼有。错误的图表解释会降低性能,即使是在模型仅从文本就能正确解决的问题上也是如此,并且准确率通常从原始图像到人工撰写的标题逐步提高。在修正标题下的恢复率对于某些模型来说很高,这将感知受阻的失败与真正的推理瓶颈区分开来。感知失败后出现的推理错误类型取决于领域:物理失败归结为计算错误,几何失败则归结为概念误用。作为核心实验之外的讨论,InternS1-mini尽管经过了大量的科学预训练并具备思考能力,但在每项任务上都低于实验中最弱的模型,其推理轨迹经常在完成前截断。
英文摘要
Multimodal LLMs report strong performance on scientific reasoning benchmarks, yet most treat perception and reasoning as a single measurable process. We introduce a five-task diagnostic experiment across physics and geometry benchmarks that isolates failures to perception, reasoning, or both. Incorrect diagram interpretation degrades performance even on problems models solve correctly from text alone, and accuracy generally rises from raw images to human-authored captions. Recovery under corrected captions is high for some models, separating perception-blocked failures from genuine reasoning bottlenecks. Which reasoning error follows a perception failure depends on domain: physics failures resolve into calculation errors, geometry into conceptual misapplication. As a discussion beyond our core experiments, InternS1-mini, despite heavy scientific pretraining and thinking capabilities, falls below the weakest model from experiments on every task, with reasoning traces frequently truncating before completion.