发表机构
Michigan State University; University of Michigan(密歇根州立大学; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MoTIF-X通过以基序为中心的层次对比学习和多模态掩码建模,统一分子图、SMILES和3D构象,实现可解释的分子表示,在OpenADMET和药物-靶点预测中表现最优。
AI 中文摘要
分子表示学习是计算机辅助药物发现的核心。分子图、SMILES字符串和3D构象提供了互补的结构信息,然而许多多模态方法独立编码这些视图,仅在后期进行对齐,限制了细粒度的跨模态交互和子结构级别的可解释性。为解决这些局限性,我们引入了MoTIF-X,一种以基序为中心的框架,使用基于图的化学基序作为多模态集成和解释的共享锚点。其第一阶段预训练通过跨原子、基序和分子尺度的层次对比学习来学习基序表示。第二阶段通过多模态掩码标记建模,将这些表示与SMILES和扭转角标记进行上下文化。在包含多种构象的药物样分子上进行预训练后,MoTIF-X在所有九个OpenADMET ExpansionRx端点上取得了最低的平均绝对误差,并在评估方法中取得了最佳整体性能。显著性分析支持其在多次检验校正后绝大多数端点-基线比较中的优势。消融研究支持了基序标记上下文化、多模态集成和两阶段预训练的互补贡献。除分子性质外,该框架扩展至药物-靶点相互作用预测,在评估基准中取得了最佳平均分类性能,并无需额外微调即可泛化至外部药物冷启动数据集。其以基序为中心的设计还实现了子结构级别的解释:较高的基序归因分数与较大的实验测量活性变化相关。总之,这些发现支持MoTIF-X作为分子建模的可迁移且可解释的框架。
英文摘要
Molecular representation learning is central to computer-aided drug discovery. Molecular graphs, SMILES strings, and 3D conformations provide complementary structural information, yet many multimodal approaches encode these views independently and align them only at a later stage, limiting fine-grained cross-modal interaction and substructure-level interpretability. To address these limitations, we introduce MoTIF-X, a motif-centered framework that uses graph-grounded chemical motifs as shared anchors for multimodal integration and interpretation. Its first pretraining stage learns motif representations through hierarchical contrastive learning across atomic, motif, and molecular scales. The second stage contextualizes these representations with SMILES and torsion-angle tokens through multimodal masked token modeling. After pretraining on drug-like molecules with multiple conformers, MoTIF-X achieved the lowest mean absolute error on all nine OpenADMET ExpansionRx endpoints and the best overall performance among the evaluated methods. Significance analyses supported its advantage in the vast majority of endpoint-baseline comparisons after multiple-testing correction. Ablation studies supported the complementary contributions of motif-token contextualization, multimodal integration, and two-stage pretraining. Beyond molecular properties, the framework extended to drug-target interaction prediction, achieving the best average classification performance across the evaluated benchmarks and generalizing to an external drug-cold-start dataset without additional fine-tuning. Its motif-centered design also enabled substructure-level interpretation: higher motif attribution scores were associated with larger experimentally measured activity shifts. Together, these findings support MoTIF-X as a transferable and interpretable framework for molecular modeling.
Comments5 figures. Supplementary material is available as an ancillary file. Code: https://github.com/Bin-Chen-Lab/Motif-X