发表机构
Zhejiang University; Zhejiang Lab(浙江大学; 之江实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM驱动科学实验中智能体的方法学幻觉问题,提出ABE-Ralph审计框架,在复现实验中实现93%稳健执行率,可检测科学失败模式,NatureBench任务表现优异。
AI 中文摘要
用于科学实验的大语言模型(LLM)智能体必须做的不仅仅是生成可执行代码:它们必须忠实地实现参考方法,设计能检验论文主张的实验,并提供支持这些主张的证据。我们发现,这类智能体常产生方法学幻觉:悄悄减少数据集或训练预算,用查找函数或神谕函数替代失效的学习或生成组件,或从资源受限的环境中得出结论——在这种环境下,方法所宣称的优势会消失。为检测这些失败,我们引入ABE-Ralph,一种锚定参考的审计框架,它将主张、协议、所需组件、基准和指标表示为结构化实验约束,通过8步工作流程指导实现,并执行定量、定性和代码级别的验证。在覆盖12个机器学习领域的30次长周期复现运行中,ABE-Ralph达到93%的稳健执行率,并识别出5种科学失败模式。在23项NatureBench发现任务中,ABE-Ralph在5项任务上的表现与最先进水平相当或超过。这些结果表明,对AI科学家的可靠评估必须评估实验设计是否忠实地检验了预期主张,以及所得证据是否支持该主张,而非将代码执行或看似合理的指标视为科学成功的证据。
英文摘要
LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfully, design experiments that test the paper's claims, and provide evidence supporting those claims. We show that agents often produce methodological hallucinations: silently reducing datasets or training budgets, replacing failed learning or generative components with lookup or oracle functions, or drawing conclusions from resource-limited settings where a method's claimed advantage disappears. To detect these failures, we introduce ABE-Ralph, a reference-anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8-step workflow, and performs quantitative, qualitative, and code-level verification. Across 30 long-horizon reproduction runs covering 12 machine learning domains, ABE-Ralph achieves a 93% robust execution rate and identifies five scientific failure modes. In 23 NatureBench discovery tasks, ABE-Ralph matches or exceeds state-of-the-art performance on 5 tasks. These results show that reliable evaluation of AI scientists must assess whether the experimental design faithfully tests the intended claim and whether the resulting evidence supports it, rather than treating code execution or plausible metrics as evidence of scientific success.
Comments20pages, 5 figures, code link: https://github.com/Flavorfish/AutoRepro