AI 中文总结
提出Sci-MMR基准,通过235个多跳任务评估多模态智能体的多步证据推理,发现答案准确率与证据恢复率差距超20%,并定位证据获取与整合两大瓶颈。
AI 中文摘要
自主研究智能体日益被期望能够检索文献、分析实验证据并生成科学假设。这些能力需要多步证据支撑的推理,即在得出结论之前逐步获取、整合并验证证据。然而,现有的多模态基准大多评估最终答案的准确性,而未揭示预测是否真正得到可追溯的科学证据的支持。我们提出了Sci-MMR,一个基于结构化论证图的多步证据支撑科学推理基准,该图将科学主张、引文支撑的知识、视觉证据和支持区域联系起来。Sci-MMR包含235个跨四个科学学科的多跳推理任务,每个任务平均包含九个图面板。评估八个前沿多模态模型后,我们发现答案准确性始终比完整证据恢复率高超过20%,揭示了一个结构性无法被仅答案评估捕获的显著差距。通过受控干预,我们识别出两个根本瓶颈。第一,证据获取:模型难以从科学图中提取完整的结构化证据,占失败原因的57.2%。虽然裁剪工具带来适度提升(+4.5个百分点),但提供金标准证据可将准确性提升高达37.0个百分点,表明模型在组装完整多区域证据方面存在困难。第二,证据整合:模型难以将可用证据转化为正确结论,占失败原因的31.8%,即使使用金标准证据,最强模型在最难任务上的准确性也仅达到69.1%。这些发现表明,当前以答案为中心的基准显著高估了多模态研究智能体的证据支撑推理能力。
英文摘要
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents