arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BioStudyBench:评估智能体在知识截止日期后的生物医学研究中的表现

BioStudyBench: Evaluating Agents on Post-Cutoff Biomedical Studies

David Li, Shaamil Karim, Christian Gensbigler

arXiv 2610.07614首次发表:更新:

发表机构

University of Waterloo; Atlas Discovery(滑铁卢大学; Atlas Discovery)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BioStudyBench通过25个长周期任务评估AI智能体在知识截止日期后仅用公开数据复现生物医学研究的能力,发现数据和工具访问使通过率平均提升47个百分点,且闭源模型优于开源模型。

AI 中文摘要

我们评估AI智能体能否仅使用公开数据复现已发表生物医学研究报告的发现。现有评估未能一致地将分析过程与先验知识或对已发表答案的检索区分开来。我们引入了BioStudyBench,一个包含25个长周期分析任务的基准,这些任务选自2026年7月至9月首次发表的研究,这些研究晚于我们评估的模型开发者报告的知识截止日期,并通过半自动筛选从404,019条PubMed记录中筛选而来。在每个任务中,智能体接收一个中性的研究问题,但不提供数据文件,因此它必须查找并下载相关的公开数据,通过仅返回截止日期之前记录的工具搜索文献,并通过数据分析报告发现。为了衡量相对于先验知识的提升,我们在有无数据和工具访问两种情况下运行每个任务。在八个模型中,数据和工具的访问使通过率相对于无数据基线平均提高了47个百分点。各种规模的开源模型落后于闭源模型,最佳开源模型通过了81.3%的任务,而最佳闭源模型通过了94.7%。

英文摘要

We evaluate whether AI agents can match the reported findings of published biomedical studies using public data. Existing evaluations do not consistently separate analysis from prior knowledge or retrieval of the published answer. We introduce BioStudyBench, a benchmark of 25 long-horizon analysis tasks drawn from studies first published between July and September 2026, after the developer-reported knowledge cutoffs of the models we evaluate, semi-automatically filtered down from 404,019 PubMed records. In each task, the agent receives a neutral research question but no data files, so it must find and download the relevant public data, search the literature through tools that return only records dated before its cutoff, and report findings through data analysis. To measure gains over prior knowledge, we run every task both with and without access to data and tools. Across eight models, access to data and tools raises the pass rate by 47 percentage points on average over the no-data baseline. Open-weight models across sizes trail closed-weight models, with the best open-weight model passing 81.3% of tasks against 94.7% for the best closed-weight model.

CommentsAccepted into AgenticLS (NeurIPS 2026 workshop)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑