发表机构
Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Trèfle通过强化学习后训练将物理约束摊销进一步流模型,以30倍吞吐量生成准确过渡态,少样本迁移到新金属并零样本泛化到更大分子,显著提升反应发现效率。
AI 中文摘要
过渡态搜索仍然是反应发现中的主要瓶颈,因为识别有效的鞍点结构需要大量昂贵的量子化学计算。生成模型可以通过从反应物和产物几何结构提出候选结构来减轻这一负担,但仅基于几何结构的监督训练并不强制满足过渡态所需的物理条件,而在采样过程中强制执行这些条件则成本高昂。我们引入了Trèfle,一种一步流模型,将物理引导的生成摊销到训练中:均值流目标替代多步采样,低成本代理替代密度泛函理论(DFT)奖励评估,以及强化学习后训练(使用强制这些条件的奖励)替代推理时的引导。Trèfle以多步基线30倍的吞吐量生成准确的过渡态。它能够少样本迁移到十种未见过的过渡金属,并零样本泛化到远大于训练中任何分子的分子。结合DFT精炼和反应路径验证,Trèfle种子搜索恢复了全部13个未见的γ-酮氢过氧化物通道,而后训练将85个未见双分子通道的单次尝试恢复率从28%提高到50%。因此,Trèfle为自动化反应发现提供了快速、可靠的种子。
英文摘要
Transition-state searches remain a major bottleneck in reaction discovery, as identifying valid saddle-point structures requires numerous expensive quantum-chemical calculations. Generative models can reduce this burden by proposing candidates from reactant and product geometries, but supervised training on geometries alone does not enforce the physical conditions required of a transition state, while enforcing them during sampling is costly. We introduce Trèfle, a one-step flow model that amortizes physics-guided generation into training: a mean-flow objective replaces multi-step sampling, a low-cost surrogate replaces density functional theory (DFT) reward evaluations, and reinforcement-learning post-training with rewards that enforce those conditions replaces inference-time guidance. Trèfle generates accurate transition states at $30\times$ the throughput of multi-step baselines. It transfers few-shot to ten unseen transition metals and generalizes zero-shot to molecules far larger than any in training. With DFT refinement and reaction-path validation, Trèfle-seeded searches recover all 13 unseen $γ$-ketohydroperoxide channels, while post-training raises per-attempt recovery from 28% to 50% across 85 unseen bimolecular channels. Trèfle thus provides fast, reliable seeds for automated reaction discovery.
Comments54 pages, 6 main figures, 7 extended data figures, 8 tables