多视图分子表示学习:层次图与上下文指纹
Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints
- Yonsei University(延世大学)
- University of Waterloo(滑铁卢大学)
- Yonsei University - Mirae Campus(延世大学未来校区)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
HiFi-Mol提出多视图分子预训练框架,结合层次图编码器与上下文指纹编码器,在MoleculeNet基准上实现平均ROC-AUC提升2.77%,验证了两种视图的互补性。
AI中文摘要:
分子性质预测需要能够从有限的标记数据泛化到结构新颖化合物的表示。现有的分子预训练方法通常依赖单一视图:基于图的方法建模原子-键拓扑,但提供的片段级监督有限,而指纹描述符编码化学模式,但通常作为固定的辅助特征使用。我们提出HiFi-Mol,一个多视图框架,在下游集成之前分别预训练层次图编码器和上下文指纹编码器。图分支使用片段感知掩码和多分辨率监督来捕获子结构感知表示,而指纹分支对七个指纹家族的活跃条目进行分词,并应用掩码语言建模来学习上下文嵌入。在微调期间,HiFi-Mol将投影的多分辨率图特征与指纹嵌入结合用于下游预测。在MoleculeNet基准上使用脚手架划分进行评估,HiFi-Mol在八个分类任务上相比最佳基线实现了平均ROC-AUC提升2.77%,同时在三个回归任务上保持竞争性能。进一步分析表明,片段感知掩码提高了图表示质量,分类结果显示了单个图和指纹变体的数据集依赖性优势,确认了这两个视图提供了互补的预测信号。
英文摘要:
Molecular property prediction requires representations that generalize from limited labeled data to structurally novel compounds. Existing molecular pretraining methods often rely on a single view: graph-based approaches model atom-bond topology but provide limited fragment-level supervision, whereas fingerprint descriptors encode chemical patterns but are typically used as fixed auxiliary features. We propose HiFi-Mol, a multi-view framework that separately pretrains a hierarchical graph encoder and a contextualized fingerprint encoder before downstream integration. The graph branch uses fragment-aware masking with multi-resolution supervision to capture substructure-aware representations, while the fingerprint branch tokenizes active entries from seven fingerprint families and applies masked language modeling to learn contextualized embeddings. During fine-tuning, HiFi-Mol combines projected multi-resolution graph features with fingerprint embeddings for downstream prediction. Evaluated on MoleculeNet benchmarks under the scaffold split, HiFi-Mol achieves a 2.77% improvement in average ROC-AUC over the best baseline across eight classification tasks while maintaining competitive performance on three regression tasks. Further analyses reveal that fragment-aware masking improves graph representation quality, and classification results demonstrate dataset-dependent strengths of the individual graph and fingerprint variants, confirming that the two views provide complementary predictive signals.