arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04915cs.AI

我们是在衡量科学智能吗?重新思考AI科学家的评估

Are We Measuring Scientific Intelligence? Rethinking the Evaluation of AI Scientists

Kate Zhang, Yuante Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出证据基础准确率指标,通过数据编辑测试AI科学家的评估,发现现有基准高估了智能体能力,并强调需衡量答案对证据的遵循程度。

中文摘要 AI 辅助

AI智能体现在能够端到端地执行数据驱动的科学分析,基准测试通过给智能体一个问题和一个数据集,并根据固定答案对其最终答案进行评分来评估它们。这些基准假设正确答案是从所提供的数据中得出的,我们将这一性质称为证据基础。然而,智能体也可以从先验知识或通过排除其他选项来达到答案,基于单次运行的分数无法区分这些情况。我们展示了如何测试这一假设,并发现它经常失败。对于每个问题,我们构建其数据文件的版本,其中答案的证据被保留、撤回或反转,使用预注册的参考统计量检查每次编辑,并在每个版本上运行相同的智能体。然后我们衡量证据基础的准确率,该准确率仅在智能体在证据被撤回时也能响应,并在证据被反转时遵循它的情况下,才给予正确答案。我们在BAISBench的18个单细胞问题和GeneBench-Pro的四个合成问题上评估了三种智能体框架和五个模型。在单细胞问题上,Claude智能体的准确率为95%,在没有任何数据的情况下正确回答了83%,但它们的证据基础准确率仅为41%。隐藏基因名称将十个基因任务中遵循反转证据的运行比例从58%提高到93%,这表明先验知识与所提供的数据竞争。基准分数和LLM评判者也可能奖励忽略已改变证据的答案。因此,衡量科学智能而非记忆,需要检查答案是否遵循证据,以及分数是否为此奖励它们。

英文摘要

AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot tell these cases apart. We show how to test this assumption and find that it often fails. For each question, we build versions of its data files in which the evidence for the answer is left intact, withdrawn or reversed, check each edit with a pre-registered reference statistic, and run the same agent on every version. We then measure evidence-grounded accuracy, which credits a correct answer only if the agent also responds when the evidence is withdrawn and follows it when it is reversed. We evaluate three agent scaffolds and five models on 18 single-cell questions from BAISBench and four synthetic problems from GeneBench-Pro. On the single-cell questions, the Claude agents are 95% accurate and answer 83% correctly without any data, but their evidence-grounded accuracy is only 41%. Hiding gene names raises the share of runs that follow reversed evidence from 58% to 93% on ten gene tasks, suggesting that prior knowledge competes with the supplied data. The benchmark score and LLM judges can also reward answers that ignore the changed evidence. Measuring scientific intelligence rather than recall therefore requires checking whether answers follow the evidence and whether scores reward them for it.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑