arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35515cs.AIcs.LGcs.SC

MechBench:AI科学智能体能否发现超越现象定律的机制?

MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?

  • Tsinghua University(清华大学)
  • Chinese Academy of Sciences(中国科学院)
  • Beijing Zhongguancun Academy(北京中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

Zihan Yu, Jiadong Zhang, Jialin Cheng, Jingtao Ding, Yong Li

AI总结:

MechBench基准通过分离现象定律与机制恢复,测试科学智能体发现机制的能力,实验显示智能体机制恢复准确率远低于现象定律,揭示显著泛化差距。

AI中文摘要:

科学发现不仅需要恢复描述可观测行为的数学定律,还需要识别产生这些行为的机制。现有的符号回归和科学智能体基准主要评估现象定律的恢复,而机制发现基本未得到测试。我们引入了MechBench,一个明确区分这两种能力的基准。每个任务由一个机制模型定义,该模型是一组结构化的、具有科学意义的关系,其联合后果蕴含一个可观测的现象定律,而智能体仅接收观测数据和科学背景。我们通过机制探针评估机制恢复,这些探针查询无法仅从现象定律推断的内部科学后果。为了减少对记忆中的教科书机制的依赖,我们通过对经典机制进行受控的、科学可解释的突变来构建不熟悉的变体,并筛选机制不可区分性以排除允许竞争性机制的模糊实例。在代表性科学智能体上的实验揭示了显著的“现象-机制恢复差距”:对于使用GPT-5.6-sol的Codex,在核心集上现象定律准确率达到35.00%,而机制准确率仅为13.75%,且在现象定律被正确恢复的案例中,机制恢复在64.29%的情况下失败。随着机制突变程度增加,差距扩大,即使提供正确的现象定律,机制恢复仍低于50%。这些结果揭示了机制推理中的显著泛化差距,并将机制发现确立为超越恢复可观测科学定律的独特挑战。

英文摘要:

Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery largely untested. We introduce MechBench, a benchmark that explicitly separates these two capabilities. Each task is defined by a mechanistic model, a structured set of scientifically meaningful relations whose joint consequences entail an observable phenomenal law, while agents receive only observational data and scientific context. We evaluate mechanism recovery through mechanism probes, which query internal scientific consequences that cannot be inferred from the phenomenal law alone. To reduce reliance on memorized textbook mechanisms, we construct unfamiliar variants through controlled, scientifically interpretable mutations of canonical mechanisms, and screen for mechanistic indistinguishability to exclude ambiguous instances admitting comparable competing mechanisms. Experiments across representative scientific agents reveal a substantial phenomenal--mechanism recovery gap: for Codex with GPT-5.6-sol, phenomenal-law accuracy reaches 35.00% on the Core-set while mechanism accuracy is only 13.75%, with mechanism recovery failing in 64.29% of cases where the phenomenal law is correctly recovered. The gap widens as mechanisms become increasingly mutated, and even providing the correct phenomenal law leaves mechanism recovery below 50%. These results reveal a substantial generalization gap in mechanistic reasoning and establish mechanism discovery as a distinct challenge beyond recovering observable scientific laws.

↑