SureRoute:面向无幻觉的逆合成自改进平台
SureRoute: Toward a Hallucination-Free Self-Improving Platform for Retrosynthesis
浏览论文内容
中文总结 AI 辅助
SureRoute通过化学验证器ChemHarness结合多模型集成与数据检索,抑制逆合成中的化学幻觉,在350个工业目标上实现74.3%的recall@1,并将top-1幻觉降至4.6%。
中文摘要 AI 辅助
包括大语言模型在内的人工智能模型正越来越多地融入科学发现工作流,但它们仍然容易产生幻觉。在实验科学中,此类错误直接转化为湿实验验证失败和资源浪费;在自改进的智能体系统中,自信的错误可能被强化而非纠正。逆合成提供了这种失败模式的代表性示例:现有模型能够生成化学上合理的路线,但无法可靠地确定哪些路线在实验上可行。我们将“化学幻觉”定义为一条看似有效但在竞争性反应位点、未解决的选择性或缺失机理支持下失败的路线,这种失败在很大程度上对Recall@K指标不可见。我们引入了SureRoute,一个以化学验证器为核心的逆合成平台,用于抑制化学幻觉。SureRoute结合了多模型集成、数据资产检索和ChemHarness(一个可执行的化学直觉引擎,用于路线验证和可靠性优先排序)。在350个真实工业目标的基准测试中,SureRoute达到了74.3%的recall@1,是七个单步模型和三个前沿大语言模型的2.2至3.5倍,同时将top-1化学幻觉降至4.6%,相对于前沿大语言模型减少了4至6倍。作为模型无关的重排序器,ChemHarness在任意骨干候选模型上将可检测的幻觉驱动至接近零。SureRoute表明,可靠的科学人工智能不仅需要强大的生成能力,还需要可执行的验证。
英文摘要
AI models, including large language models, are increasingly integrated into scientific discovery workflows, yet they remain prone to hallucination. In experimental sciences, such errors translate directly into failed wet-lab validations and wasted resources; in self-improving agentic systems, confident errors risk being reinforced rather than corrected. Retrosynthesis provides a representative example of this failure mode: existing models can generate chemically plausible routes, but cannot reliably determine which routes are experimentally feasible. We define \textbf{Chemical Hallucination} as a route that appears valid yet fails under competing reactive sites, unresolved selectivity, or missing mechanistic support, a failure largely invisible to the Recall@$K$ metric. We introduce \textbf{SureRoute}, a chemical verifier-anchored retrosynthesis platform that suppresses Chemical Hallucination. SureRoute combines a multi-model ensemble, data asset retrieval, and \textbf{ChemHarness}, an executable chemical intuition engine for route verification and reliability-first ranking. On a benchmark of 350 real-world industrial targets, SureRoute reaches 74.3\% recall@1, 2.2--3.5$\times$ that of seven single-step models and three frontier LLMs, while cutting top-1 Chemical Hallucination to 4.6\%, a 4--6$\times$ reduction relative to frontier LLMs. As a model-agnostic reranker, ChemHarness drives detectable hallucination toward near-zero across arbitrary backbone candidates. SureRoute shows that reliable scientific AI requires not only strong generation, but executable verification.
发表机构
- XtalPi Inc(晶泰科技)
机构由 AI 辅助整理,请以论文原文为准。