发表机构
CSIR- Central Scientific Instruments Organisation; Academy of Scientific and Innovative Research (AcSIR); Monell Chemical Senses Center; University of Pennsylvania(科学与工业研究理事会中央科学仪器组织; 科学与创新研究学院; 莫奈尔化学感官中心; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GraphNOSE提出一种图变换器框架,从SMILES字符串预测气味描述符,以少六倍参数超越基线,AUROC达84%,并利用XAI识别关键分子特征。
AI 中文摘要
从分子结构预测嗅觉品质是化学信息学中的一个开放问题。尽管线性模型可以将分子特征与气味描述符关联起来,但在外推到新颖的化学骨架、极端分子量或复杂气味混合物时,它们常常失效。为解决这一问题,我们提出了GraphNOSE,一个开源的图变换器框架,它从简化分子输入线性输入规范(SMILES)字符串预测单分子和二元混合物的多标签气味描述符。通过在基于变换器的图架构中集成位置编码和结构编码,GraphNOSE以比标准图神经网络(GNN)基线少六倍的参数实现了强劲性能,同时平均ROC曲线下面积(AUROC)比线性模型、分子语言模型嵌入、分子指纹和基线GNN高出4.52%(p < 0.01)。GraphNOSE在分布外化合物(OODs)上实现了84%的AUROC,超过了当前嗅觉OOD的最先进GNN(Open-POM:81%,p < 0.001),并识别了线性模型经验性失败的条件。最后,我们应用可解释人工智能(XAI)方法来识别哪些子结构和分子特征驱动气味预测,所得见解与化学直觉一致,并基于模型学到的表征。这些结果共同确立了GraphNOSE作为一种可扩展且可解释的嗅觉预测架构,能够泛化到当前感知数据库中代表性不足的结构上不同的化合物。
英文摘要
Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. Although linear models can link molecular features to odor descriptors, they often fail when extrapolating to novel chemical scaffolds, extreme molecular weights, or complex odor mixtures. To address this, we introduce GraphNOSE, an open-source graph transformer framework that predicts multi-label odor descriptors from simplified molecular-input line-entry system (SMILES) strings for single molecules and binary mixtures. By integrating positional and structural encodings within a transformer-based graph architecture, GraphNOSE achieves strong performance with six times fewer parameters than standard graph neural network (GNN) baseline while consistently outperforming linear models, molecular language model embeddings, molecular fingerprints, and baseline GNNs by an average area under the ROC curve (AUROC) margin of 4.52% (p < 0.01). GraphNOSE achieves an AUROC of 84% on out-of-distribution compounds (OODs). This exceeds the current state-of-the-art GNN for OOD in olfaction (Open-POM: 81%, p < 0.001), and identifies conditions under which linear models empirically fail. Finally, we apply XAI (explainable AI) methods to identify which substructures and molecular features drive odor predictions, yielding insights consistent with chemical intuition and grounded in the model's learned representations. Together, these results establish GraphNOSE as a scalable and interpretable architecture for olfactory prediction that generalizes to structurally distinct compounds underrepresented in current perceptual databases.