arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BixBench3:面向研究级计算生物学任务的AI智能体基准测试

BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks

Zane Koch, Asmamaw T. Wassie, Javier Valdes-Aleman, Jason Lee, Michaela M. Hinks, Samuel G. Rodriques, Andrew D. White, Jon M. Laurent

arXiv 2608.25286首次发表:更新:

AI 中文总结

研究人员推出BixBench3基准测试,评估AI智能体完成研究级计算生物学任务的能力,发现前沿模型表现差异大,且智能体在大数据集、多步骤任务中性能更差,高分智能体更高效。

AI 中文摘要

人工智能(AI)有望通过自动化计算分析加速生物学研究,但AI智能体完成完整研究级规模计算生物学任务的能力尚未得到系统评估。本文介绍BixBench3,这一基准测试用于衡量AI智能体从原始生物数据处理到得出科学结果的能力。我们设计BixBench3任务,使其模拟科学家将工作委托给智能体的过程:科学家选定研究问题和高级方法,再将所有分析的执行委托给智能体。在每个任务中,智能体接收研究目标、方法指导以及来自已发表科学研究的原始数据,必须执行一系列分析以达成研究目标。这些分析产生的数据产物(如峰调用矩阵或差异表达表)会与原始研究中生成并报告的对应产物进行程序化评分。在涵盖138种独特产物生成的20项BixBench3任务中,我们发现13个前沿模型的得分范围从Gemini 3.1 Flash Lite的0.00到GPT 5.6 Sol的0.48不等。智能体在原始数据集较大的任务上表现更差(<100GB的任务得分为0.36,>100GB的任务得分为0.10),且在需要更多连续步骤的分析中表现更差(1-2步得分为0.36,3步及以上得分为0.24)。平均而言,智能体完成每项任务耗时6.8小时,使用1.02亿token,成本为43美元;最长尝试耗时24小时,使用10.7亿token,成本为525美元。值得注意的是,得分最高的智能体比性能较差的选项使用更少的token且成本更低。这些结果表明,大型语言模型(LLM)在以下三方面的能力存在显著差异:(1)连贯执行多个连续分析步骤;(2)管理大量原始数据;(3)跨科学领域开展工作。

英文摘要

Artificial intelligence (AI) promises to accelerate biological research by automating computational analyses. Yet the ability of AI agents to execute on computational biology at the scale of complete research studies has not been systematically evaluated. Here we introduce BixBench3, a benchmark that measures the capacity of AI agents to process raw biological data through to scientific results. We designed BixBench3 tasks to mirror the delegation of work from a scientist to an agent: the scientist chooses the research question and high-level methods, then delegates implementation of all analyses to the agent. In each task, an agent receives a research objective, methodological guidance, and raw data derived from a published scientific study, and must execute a sequence of analyses to achieve the research objective. The data artifacts resulting from these analyses, such as peak call matrices or differential expression tables, are programmatically graded against the corresponding artifacts generated and reported in the original study. Across 20 BixBench3 tasks encompassing the generation of 138 unique artifacts, we find that 13 frontier models achieve scores ranging from 0.00 for Gemini 3.1 Flash Lite to 0.48 for GPT 5.6 Sol. Agents perform worse on tasks with larger raw datasets (0.36 on tasks with <100 GB versus 0.10 on tasks with >100 GB) and on analyses requiring more sequential steps (0.36 at 1-2 steps vs 0.24 at 3+). On average, agents use 6.8 hours, 102 million tokens, and \$43 to complete each task, with the longest attempts consuming 24 hours, 1.07 billion tokens, and \$525. Notably, the highest-scoring agents used fewer tokens and were cheaper than less performant options. These results reveal that LLMs vary substantially in their ability to (1) execute multiple sequential analysis steps coherently, (2) manage large quantities of raw data, and (3) work across scientific domains.

Comments28 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑