arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

onepot-Bench 0:面向实验室实际场景的计算机化学基准测试

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko

arXiv 2608.02595首次发表:更新:

AI 中文总结

本研究推出专有基准测试onepot-Bench 0,通过三个互补评估衡量语言模型在湿实验室合成化学场景下的基础能力、可靠性等,弥补现有评估的不足。

AI 中文摘要

语言模型在实验室科学中发挥着日益重要的作用,可完成实验规划、执行及事后分析等任务。然而,精准衡量其能力颇具难度,因为科学能力需兼具问题解决技能与领域特定直觉。现有评估极少衡量其在物理实验室中做出可靠决策所需的能力,且常依赖可能已出现在模型训练语料中的公开数据。我们推出onepot-Bench 0,这是一个专有基准测试套件,用于评估语言模型在与湿实验室执行相关的合成化学能力。onepot-Bench 0包含三个互补的评估:ChemAbacus衡量无工具的化学信息学素养与数值推理能力;SynthRefusal表征针对各类良性、受控及设计药物靶点的安全性与弃权(不执行)行为;SynthBench使用我们实验室生成的私有实验数据评估反应结果预测与催化剂选择。这些评估共同探究了基础能力、可靠性及更深入的知识,而这些均是实验室中可靠表现所需的技能。

英文摘要

Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution. onepot-Bench 0 comprises three complementary evaluations: ChemAbacus measures tool-free cheminformatics literacy and numerical reasoning; SynthRefusal characterizes safety and refusal behavior across a variety of benign, controlled, and designer-drug targets; and SynthBench evaluates reaction-outcome prediction and catalyst selection using private experimental data generated in our laboratory. Together, these evaluations probe basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑