发表机构
École Polytechnique Fédérale de Lausanne (EPFL); Laboratory for Chemical Technology, Ghent University; NCCR Catalysis(洛桑联邦理工学院; 根特大学化学技术实验室; 瑞士国家研究能力中心催化项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出多智能体LLM框架,通过验证循环自动生成反应规则,将标准分类从68类扩展到14,073类,并实现97.7%的未知反应分类准确率。
AI 中文摘要
计算机辅助合成规划使用大型反应规则库将目标分子分解为可访问的前体,每条规则为每个转化赋予确定性的、可解释的标签。但化学领域呈长尾分布,手动编码难以处理,现有工具依赖固定规则集,无法适应新化学。本文提出一个全自动流水线,其中多智能体框架的大语言模型(LLM)对665,901个美国专利反应进行分类并编写规则本身,每条规则在验证循环中生成,该循环对语料库进行测试。它将标准分类从68类扩展到14,073类,无需人工筛选。通过轻量级指纹分类器,它对97.7%的未见反应进行分类,与领先的专有分类器性能相当,同时更精细地解析化学,并按需扩展到训练分布之外的化学。结果是一个活的反应性数据库,以及将生成模型转化为可靠的、自扩展的符号系统的一般途径。
英文摘要
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encoding intractable, and existing tools rely on fixed rulesets that cannot adapt to new chemistries. Here we present a fully automated pipeline in which a multi-agent framework of large language models (LLMs) classifies reactions and writes the rules themselves across 665,901 US patent reactions, generating each rule under a verification loop that tests it against the corpus. It expands a standard taxonomy from 68 to 14,073 classes without human curation. With a lightweight fingerprint classifier, it classifies 97.7\% of unseen reactions, matching a leading proprietary classifier while resolving chemistry more finely and extending on demand to chemistry outside its training distribution. The result is a living reactivity database and a general route to turning generative models into reliable, self-expanding symbolic systems.