发表机构
Benchling(本奇林公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出BenchBench-Protocol基准测试,包含149个源自真实实验修改的湿实验室方案推理与修改任务,评估9个模型后发现Claude Opus 5得分最高,该基准测试可用于评估模型在生命科学常规任务中的表现。
AI 中文摘要
我们推出BenchBench-Protocol,这是一个针对大型语言模型的基准测试,包含149个方案修改任务,这些任务源自科学家在实际实验工作中对已发表方案所做的修改。将已发表的方案调整以适配新实验是湿实验室科学家的常规任务,正确的修改需要考虑先前的选择和下游步骤。近期生命科学基准测试已转向开放式、基于评分规则评分的任务,但这些任务通常由专家提出,而非从真实世界的修改中重构。BenchBench-Protocol的任务源自已发表方案与科学家修改后的版本之间的差异,这为查询内容和正确响应的加权评分规则要素提供了基础。该基准测试涵盖湿实验室生物学9个领域的96个源方案,仅包含经领域专家审核后评为高度优质的任务。我们评估了9个封闭和开放模型;Claude Opus 5以59.2%的归一化评分规则得分位居榜首,其他模型的得分在34.1%至47.1%之间,且在采用10次尝试中的最佳结果时,该基准测试仍未饱和。随着模型在生命科学研究中日益发挥帮助作用,在常规湿实验室任务上对其进行评估的重要性也相应提升。我们推出BenchBench-Protocol,作为对湿实验室推理的 grounded 评估,以及利用真实世界实验构建基准测试任务的实用性的证据。
英文摘要
We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recovered from modifications that scientists made to published protocols during real experimental work. Adapting a published protocol to a new experiment is a routine task for a wet-lab scientist, and a correct modification requires accounting for prior choices and downstream steps. Recent life-science benchmarks have moved toward open-ended, rubric-graded tasks, but tasks are typically elicited from experts rather than reconstructed from real-world modifications. BenchBench-Protocol tasks are derived from differences between a published protocol and a version a scientist modified, which provides the basis for the query and the weighted rubric elements for a correct response. The benchmark draws from 96 source protocols across nine domains of wet-lab biology and only includes tasks rated highly after review by domain experts. We evaluate nine closed and open models; Claude Opus 5 scores highest at 59.2% normalized rubric score, with other models between 34.1% and 47.1%, and the benchmark remains unsaturated when taking the best of ten attempts. As models are increasingly helpful in life-sciences research, evaluating them on routine wet-lab tasks becomes correspondingly important. We present BenchBench-Protocol as both a grounded assessment of wet-lab reasoning and evidence for the utility of real-world experiments to construct benchmark tasks.
Comments21 pages, 15 figures