arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FREA:面向反应可行性验证的多源专家基准

FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification

Botao Yu, Bo Zhou, Daniel Adu-Ampratwum, Frazier N. Baker, Ziru Chen, Reza Averly, Ye Liu, Wenhao Gao, Xia Ning, Huan Sun

arXiv 2610.06614首次发表:更新:

发表机构

The Ohio State University; University of Illinois Chicago; University of Pennsylvania(俄亥俄州立大学; 伊利诺伊大学芝加哥分校; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FREA基准评估反应可行性验证器,发现无验证器在所有来源上领先,前向模型在逆合成提议上最佳但拒绝可行编辑,负监督提升有限,需跨来源评估并设计迁移负候选。

AI 中文摘要

随着生成模型和AI智能体以超出专家审查的规模提出化学反应,可行性验证器决定哪些提议进入合成规划。但它们的决策在不同类型的候选反应上与化学家是否一致?我们引入了FREA,一个由751个反应组成的基准,这些反应由专家化学家根据明确的可行性标准进行标注,来源包括逆合成模型提议、零产率实验记录、大语言模型(LLMs)的编辑以及五种负候选生成方法。我们的评估发现,没有验证器在所有来源上均领先:仅给定标准的LLMs与专用验证器相当,而前向模型在逆合成提议上表现最佳,但在评估的操作点上拒绝了大多数记录反应的可行的编辑。超越总体得分,在区分不可行的替代断键与可行的生成候选时,两个前向模型的表现均低于随机水平。为了研究负监督是否能解决这些弱点,我们还发布了一个包含超过1400万条记录反应和生成负候选的语料库。在匹配的训练比较中,将生成的负候选混合加入前向训练可提高各来源的平均AUROC,但这些增益并未扩展到逆合成提议。进一步改变生成方法表明,生成候选上的最大增益与更差的提议筛选同时发生。这些发现激励了在不同来源上针对专家评估验证器,并设计负候选以迁移到模型提议。

英文摘要

As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce FREA, a benchmark of 751 reactions labeled by expert chemists under an explicit feasibility criterion, drawn from retrosynthesis model proposals, zero-yield experimental records, edits by large language models (LLMs), and five negative candidate generation methods. Our evaluation finds that no verifier leads across all sources: LLMs given only the criterion are competitive with dedicated verifiers, while forward models perform best on retrosynthesis proposals but reject most feasible edits of recorded reactions at the evaluated operating points. Looking beyond aggregate scores, both forward models perform below chance when separating infeasible alternative disconnections from feasible generated candidates. To study whether negative supervision addresses these weaknesses, we also release a corpus of over 14 million recorded reactions and generated negative candidates. In matched training comparisons, adding a mixture of generated negatives to forward training raises mean AUROC across sources, but these gains do not extend to retrosynthesis proposals. Varying the generation method further shows that the largest gain on generated candidates coincides with worse proposal screening. These findings motivate evaluating verifiers against experts across sources and designing negatives for transfer to model proposals.

CommentsOngoing work

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑