arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向反应产率预测的角色感知摩根指纹

Role-Aware Morgan Fingerprints for Reaction Yield Prediction

Chinmay Mirji, Prashant Shekhar, Foram Madiyar, Hao Peng

arXiv 2609.22167首次发表:更新:

发表机构

Embry-Riddle Aeronautical University; Bethune Cookman University(安柏瑞德航空大学; 贝休恩-库克曼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于角色感知摩根指纹的MFP方法,通过角色聚合和差分特征构建反应描述符,在Suzuki-Miyaura和Buchwald-Hartwig基准上实现高准确率(R2达0.969)且训练速度快一个数量级,为反应产率预测提供高效基线。

AI 中文摘要

从分子结构和反应背景预测反应产率可以减少合成化学中的实验试错,并加速条件筛选。近期用于该任务的方法使用学习表示,如图神经网络或对反应SMILES(简化分子输入线性输入系统)的Transformer编码器,但这些方法带来繁重的预处理开销,并且在输入格式不一致时可能失效。我们提出MFP,一种基于角色感知摩根指纹的反应产率预测方法,其中对每个反应组分计算基于计数的环形指纹,按化学角色(反应物、试剂、产物)聚合,并与变换敏感的差分特征组合成固定长度的反应描述符,输入前馈神经回归器。我们在Suzuki-Miyaura和Buchwald-Hartwig基准上,使用共享的预处理和评估协议,将MFP与最先进方法如YieldBERT(有和没有数据增强)和GNAN(图神经网络)进行比较。MFP在Suzuki-Miyaura上达到R2=0.878,在Buchwald-Hartwig上达到R2=0.969,同时训练速度比基于图或Transformer的替代方法快一个数量级。正式复杂度分析确认MFP将所有表示成本折叠到一次性预处理步骤中,消除了图方法所携带的每轮消息传递开销。对指纹半径和折叠向量长度的消融研究表明,在nBits=2048时,半径-2表示在两个数据集上提供了准确性、速度和跨分割稳定性的最佳平衡。这些结果确立了MFP作为反应产率预测的有效、可复现且高效的基线。

英文摘要

Predicting reaction yield from molecular structure and reaction context can cut experimental trial-and-error and speed up condition screening in synthetic chemistry. Recent methods for this task use learned representations such as graph neural networks or Transformer encoders over reaction SMILES (Simplified Molecular Input Line Entry System), but these approaches carry heavy preprocessing overhead and can break when input formatting is inconsistent. We propose MFP, a reaction yield prediction method built on role-aware Morgan fingerprints where count-based circular fingerprints are computed for each reaction component, aggregated by chemical role (reactant, reagent, product), and combined with transformation-sensitive difference features into a fixed-length reaction descriptor fed to a feed-forward neural regressor. We test MFP against state of the art methods such as YieldBERT (with and without data augmentation) and GNAN (graph neural network) on the Suzuki-Miyaura and Buchwald-Hartwig benchmarks using a shared preprocessing and evaluation protocol. MFP reaches R2 = 0.878 on Suzuki-Miyaura and R2 = 0.969 on Buchwald-Hartwig while training an order of magnitude faster than graph- or Transformer-based alternatives. A formal complexity analysis confirms that MFP folds all representation cost into a one-time preprocessing step, removing the per-epoch message-passing overhead that graph methods carry. An ablation over fingerprint radius and folded vector length shows that radius-2 representations at nBits =2048 give the best balance of accuracy, speed, and cross-split stability on both datasets. These results establish MFP as an effective, reproducible, and efficient baseline for reaction yield prediction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑