arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从科学观测到机制:AI 科学家假设生成的基准测试

From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists

Xiaxun Xie, Qingqing Long, Meng Xiao, Wei Ju, Yuanchun Zhou, Xuezhi Wang, Hengshu Zhu

arXiv 2610.05197首次发表:更新:

发表机构

Computer Network Information Center, Chinese Academy of Sciences; National University of Singapore; Sichuan University(中国科学院计算机网络信息中心; 新加坡国立大学; 四川大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 MechHypoBench 基准,用于评估 AI 智能体从经验数据生成机制性假设的能力,实验发现其生成假设与真实机制存在显著差距。

AI 中文摘要

数据驱动的机制性假设对科学发现至关重要,因为它们解释了潜在过程如何产生观测到的现象。AI 智能体和 AI 科学家越来越多地支持科学数据分析。然而,他们从经验发现中提炼机制性假设的能力仍未得到充分检验。为填补这一空白,我们引入了 MechHypoBench,这是首个用于评估 AI 智能体和 AI 科学家能否从经验数据中生成此类假设的基准。它结合了来自 14 个科学领域的论文推导机制和包含 1798 万条记录的真实世界数据集。该构建保留了经验数据的观测复杂性,同时提供了指定的潜在机制。智能体分析观测结果并提出开放式假设。我们开发了一个评估框架,通过保留条件下的后果来评估开放式机制性假设。针对通用智能体和 AI 科学家的实验表明,生成的假设与潜在机制之间存在显著差距。

英文摘要

Data-driven mechanistic hypotheses are essential to scientific discovery because they explain how underlying processes produce observed phenomena. AI agents and AI scientists increasingly support scientific data analysis. However, their ability to turn empirical findings into mechanistic hypotheses remains insufficiently examined. To address this gap, we introduce MechHypoBench, the first benchmark for evaluating whether AI agents and AI scientists can generate such hypotheses from empirical data. It combines paper-derived mechanisms from 14 scientific fields with real-world datasets containing 17.98 million records. The construction retains the observational complexity of empirical data while providing a specified underlying mechanism. Agents analyze the observations and propose open-form hypotheses. We develop an evaluation framework that assesses open-form mechanistic hypotheses through their consequences under withheld conditions. Experiments with general agents and AI scientists reveal a substantial gap between generated hypotheses and the underlying mechanisms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑