arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoSciBench:用于评估科学智能体的自主基准生成

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

Dongki Kim, Namkyeong Lee, Surag Nair, Carl Edwards, Xiner Li, Edward De Brouwer, Jenna Lynn Collier, Sung Ju Hwang, Gabriele Scalia, Ehsan Hajiramezanali

arXiv 2610.05140首次发表:更新:

发表机构

KAIST; Genentech(韩国科学技术院; 基因泰克公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AutoSciBench框架,通过自动生成和迭代精炼科学任务基准,降低求解准确率并提升任务质量,实现科学智能体评估的动态适应。

AI 中文摘要

随着智能体的快速发展,现有基准可能变得饱和,从而限制其区分能力并揭示剩余失败模式的能力。尤其在科学领域,构建和更新基准需要大量时间、人力和领域专业知识,使得评估难以与智能体能力的进步保持同步。我们通过研究科学智能体基准是否能够自动生成并随着智能体能力的演化而迭代调整来应对这一挑战。我们引入了AutoSciBench,一个将每个任务表示为高层概念的框架,该概念指定科学领域、数据模态和所需推理方法,同时包含一个低层配方,规定问题、环境和真实答案的构建与验证方式。智能体尝试解决每个任务,产生求解轨迹和相应的评判反馈,AutoSciBench利用这些反馈来修订配方或概念,关闭观察到的捷径,并将任务转向原始数据复查、中间结果解释和证据整合。从已完成精炼轨迹中提取的经验进一步指导新概念的生成,使早期任务精炼的教训能够为后续基准构建提供信息。从现有基准出发,我们在计算生物学、材料科学和临床影像领域评估了AutoSciBench。生成的基准相对于人工策划的基准,在计算生物学和材料科学中分别将平均求解准确率降低了22.4和25.5个百分点,而生成的任务在所有三个领域中获得更高的平均质量评分,表明科学智能体评估能够随着智能体能力的进步而适应。

英文摘要

As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain expertise, making it difficult to keep evaluation aligned with advances in agent capabilities. We address this challenge by investigating whether scientific-agent benchmarks can be automatically generated and iteratively adapted as agent capabilities evolve. We introduce AutoSciBench, a framework that represents each task as a high-level concept specifying the scientific domain, data modality, and required reasoning approach, together with a low-level recipe specifying how the question, environment, and ground-truth answer are constructed and verified. Agents attempt to solve each task, producing solver trajectories and corresponding judge feedback which AutoSciBench uses to revise the recipe or concept, closing observed shortcuts and shifting tasks toward raw-data re-examination, interpretation of intermediate results, and evidence integration. Experience distilled from completed refinement trajectories further guides new concept generation, allowing lessons from earlier task refinement to inform subsequent benchmark construction. Starting from existing benchmarks, we evaluate AutoSciBench across computational biology, materials science, and clinical imaging. Generated benchmarks reduce average solver accuracy by 22.4 and 25.5 percentage points relative to the human-curated benchmarks in computational biology and materials science, respectively, while generated tasks receive higher average quality ratings across all three domains, suggesting that scientific-agent evaluation can adapt as agent capabilities advance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑