发表机构
School of Cyber Science and Engineering, Wuhan University; School of Computer Science, Wuhan University; College of Computer Science and Technology, Zhejiang University; College of Chemistry and Molecular Sciences, Wuhan University(武汉大学网络安全学院; 武汉大学计算机学院; 浙江大学计算机科学与技术学院; 武汉大学化学与分子科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RetroMPA是一种与模型无关的即插即用化学过滤器,通过注入化学知识增强逆合成预测,在USPTO-50K和USPTO-Full数据集上提升了多种模型的准确率,湿实验验证了其实际效用。
AI 中文摘要
逆合成是药物发现与有机合成的基石。尽管数据驱动的深度学习模型已取得显著进展,但它们从大规模数据集中自主学习反应模式时,较少将已确立的化学知识作为先验进行整合。为解决这一局限,我们提出RetroMPA,一种分子属性感知的事后增强模块,可将化学知识注入逆合成流程。RetroMPA并非独立的SMILES序列生成器,而是一种适用范围广泛、与模型无关的化学过滤器,旨在重新校准并优化现有算法的预测路径。这一即插即用框架可无缝集成多种数据驱动的逆合成方法,无需修改模型架构或进行资源密集型的重新训练,即可增强输出效果。通过利用属性感知的潜在嵌入空间,RetroMPA在USPTO-50K数据集上,使8种代表性逆合成模型的Top-1准确率平均提升5.50%;此外,我们在大规模USPTO-Full数据集上验证了其可扩展性,在基于模板和无模板的架构上均实现约2.03%的平均提升。湿实验为该框架的实际实用性提供了初步支持,这些合成实验证实了经典反应范式(即Suzuki-Miyaura偶联、布赫雷尔反应、傅-克酰基化)的可行且此前未被报道的底物组合,表明RetroMPA的作用不止于数据拟合。代码已在该https URL开源。
英文摘要
Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progress, they autonomously learn reaction patterns from extensive datasets with limited integration of established chemical knowledge as priors. To address this limitation, we introduce RetroMPA, a molecular property-aware, post-hoc enhancement module that injects chemical knowledge into the retrosynthesis pipeline. Rather than functioning as an independent SMILES sequence generator, RetroMPA is a broadly applicable, model-agnostic chemical filter designed to recalibrate and optimize the predictive pathways of existing algorithms. This plug-and-play framework integrates seamlessly with a range of data-driven retrosynthesis methods, enhancing outputs without modifying model architecture or requiring resource-intensive retraining. By leveraging a property-aware latent embedding space, RetroMPA consistently improves top-1 accuracy across eight representative retrosynthesis models by an average of 5.50% on USPTO-50K. Furthermore, we validate its scalability on the large-scale USPTO-Full dataset, achieving an average improvement of about 2.03% across both template-based and template-free architectures. Wet-lab experiments provide preliminary support for the practical utility of the framework. These syntheses confirmed viable, previously unreported substrate combinations for classic reaction paradigms---specifically, Suzuki-Miyaura coupling, Bucherer reaction, and Friedel-Crafts acylation---suggesting that RetroMPA can operate beyond mere data fitting. The code is open-sourced at https://github.com/MengzhouLu/RetroMPA.
CommentsAccepted for publication in Journal of Chemical Information and Modeling