发表机构
HKUST; Alibaba Group(香港科技大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SciRIGOR通过重建证据图并评分完整论断-支持路径,评估科学编码智能体在六个领域100个案例中的表现,发现内部一致性不等于科学正确性。
AI 中文摘要
科学编码智能体产生相互依存的代码、结果、图表和论断,然而仅评估最终输出并不能确定其结论是否得到科学支持。我们构建了基于证据的多模态科学分析,要求智能体生成可执行的分析以及由同一运行中的结果和可视化支持的论断。我们引入SciRIGOR,一个评估框架和基准,包含来自六个领域和17个子领域的科学文章中的100个案例。该框架重构类型化证据图,将工件保真度与关系有效性分离,并评分完整的论断-支持路径,同时定位最早的不支持关系。基于来源的替代路径可容纳科学上等效的分析和可视化。我们评估了11种智能体/模型配置。在完整基准运行中,论断与忠实和不忠实结果的一致性率几乎相同(91.8%对91.0%)。然而,在软证据链评分上,没有系统超过62.6%,严格全链成功率最高为18.0%。这些发现表明,内部一致性并不能确立科学正确性:评估必须验证从数据到论断的完整路径上的支持。
英文摘要
Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supported by results and visualizations from the same run. We introduce SciRIGOR, an evaluation framework and benchmark comprising 100 cases from scientific articles across six domains and 17 subfields. The framework reconstructs typed evidence graphs, separates artifact fidelity from relational validity, and scores complete claim-support paths while localizing the earliest unsupported relation. Source-grounded alternative paths accommodate scientifically equivalent analyses and visualizations. We evaluate 11 agent/model configurations. On full-benchmark runs, claims agree with faithful and unfaithful results at nearly identical rates (91.8% versus 91.0%). Yet no system exceeds 62.6% on the soft evidence-chain score or 18.0% strict whole-chain success. These findings show that internal coherence does not establish scientific correctness: evaluation must verify support along the complete data-to-claim path.