arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04261q-bio.QMcs.LGstat.ML

基于LeJEPA的分子图编码器自监督预训练

Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA

Michał Kulczykowski, Rafał Łabędzki

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将LeJEPA适配至分子图,发现其预训练虽能提升分子图编码器表示性能,但微调增益弱且依赖划分,嵌入与Morgan指纹结合可显著提升ogbg-molhiv预测性能。

中文摘要 AI 辅助

自监督预训练已在语言和视觉领域取得变革性进展,但其在分子图神经网络中的价值仍存在争议。本研究探究在大规模未标注语料库上进行预训练是否能提升分子性质预测性能。我们将LeJEPA(一种由草图各向同性高斯正则化(SIGReg)正则化的无预测器联合嵌入预测架构)适配至分子图,采用多种子自举协议,在Wong等人[1]的抗生素活性数据集和ogbg-molhiv数据集上评估GPS及Chemprop风格的D-MPNN编码器。预训练可改进学习到的表示,但无法稳健提升微调效果:在两个任务上,对预训练嵌入的冻结探针(frozen probe)均优于随机初始化(ogbg-molhiv的ROC-AUC为0.788,对比随机初始化的0.665,提升0.123),达到已发表的自监督水平,但未转化为微调增益;在抗生素支架划分中,单一典型划分的AUPRC提升显著(delta AUPRC +0.041,p=0.010),但在五个划分合并后效果消失(合并后delta +0.013,p=0.095);在随机划分、ogbg-molhiv及D-MPNN上,微调效果为零。不过该表示优势仍可恢复:嵌入在约16-32个有效维度时达到饱和,而Morgan指纹则提升至1024位;在匹配维度下,指纹在验证集表现领先(128维时为0.799,对比嵌入的0.782),但在移位测试支架上落后(0.759 vs 0.788);将嵌入截断并与1024位Morgan指纹结合后,ogbg-molhiv的ROC-AUC从0.805提升至0.832(delta +0.027;95%置信区间[+0.003, +0.054];p=0.014),而未预训练的编码器无增益(delta -0.003)。综上,预训练提供的互补信息最好通过特征级结合实现,而微调增益较弱且依赖划分。

英文摘要

Self-supervised pretraining has transformed language and vision, but its value for molecular graph neural networks remains contested. We ask whether pretraining on a large unlabelled corpus improves molecular property prediction. We adapt LeJEPA, a predictor-free joint-embedding predictive architecture regularised by Sketched Isotropic Gaussian Regularisation (SIGReg), to molecular graphs, evaluating GPS and Chemprop-style D-MPNN encoders on the Wong et al. [1] antibiotic-activity dataset and ogbg-molhiv using a multi-seed, bootstrap-based protocol. Pretraining improves learned representations but does not robustly improve finetuning. A frozen probe on pretrained embeddings exceeds random initialisation on both tasks (ogbg-molhiv ROC-AUC 0.788 vs 0.665; +0.123), reaching the published self-supervised band, but this does not translate into finetuning gains. On the antibiotic scaffold split, a canonical partition is significant (delta AUPRC +0.041, p = 0.010), but the effect vanishes across five partitions (pooled +0.013, p = 0.095). Finetuning is null on the random split, ogbg-molhiv, and D-MPNN. The representational edge is nevertheless recoverable. Embeddings saturate at ~16-32 effective dimensions, whereas Morgan fingerprints improve to 1024 bits. At matched dimensionality, fingerprints lead validation (0.799 vs 0.782 at 128 dimensions) but trail shifted test scaffolds (0.759 vs 0.788). Truncating embeddings and combining them with a 1024-bit Morgan fingerprint raises ogbg-molhiv ROC-AUC from 0.805 to 0.832 (delta +0.027; 95% CI [+0.003, +0.054]; p = 0.014); an untrained encoder gains nothing (delta -0.003). Thus, pretraining supplies complementary information best realised through feature-level combination, while finetuning gains are weak and partition-dependent.

发表机构

  • SGH Warsaw School of Economics(华沙经济学院)

机构由 AI 辅助整理,请以论文原文为准。

↑