AI 中文总结
研究针对分子构象集合预测性质的问题,提出EnsembleEGNN模型,先以EGNN层编码各构象,再用集合注意力块汇集表示,经多任务自监督预训练,在环肽数据集上取得良好效果,联合BERT编码器进一步提升预测性能。
AI 中文摘要
从结构预测分子性质时通常使用单个代表性构象,而许多分子在溶液中以构象集合形式存在。我们引入了EnsembleEGNN,这是一种分子集合基础模型,通过共享的等变图神经网络(EGNN)层对每个构象进行编码,然后用集合注意力块汇集所得构象表示来对集合进行编码。我们在环肽集合数据集CREMP上使用多任务自监督目标(结合掩码令牌恢复、噪声坐标重建和成对距离重建)对模型进行预训练。在CREMP - CycPeptMPDB数据集上,从头开始训练EnsembleEGNN完全失败(\(R^2 = 0.005\))。然而,预训练模型达到了\(R^2 = 0.477\)和皮尔逊\(r = 0.699\),优于仅序列的BERT基线(\(R^2 = 0.439\),皮尔逊\(r = 0.667\))。当EnsembleEGNN与BERT序列编码器端到端联合训练时,混合模型进一步提高到\(R^2 = 0.538\)和皮尔逊\(r = 0.737\)。这些结果表明,将构象集合编码为单个热力学信息嵌入可改善环肽性质预测。
英文摘要
Molecular graph encoding often relies on a single, static structure, ignoring the thermodynamic ensemble of molecules that are present in solution. Here, we introduce EnsembleEGNN, a foundation model that encodes structural ensembles by processing individual conformers through shared equivariant graph neural network layers, pooled with a set attention block, to make property predictions from the whole ensemble. Pretrained on the CREMP cyclic peptide dataset using multi-task self-supervision, the model is trained to encode the conformational variability of each molecule. When predicting membrane permeability from the CycPeptMPDB benchmark, EnsembleEGNN achieves an $R^2$ of $0.477$ under random cross-validation, outperforming a sequence-only BERT baseline ($R^2=0.439$). This representation advantage persists under rigorous out-of-distribution Butina splits ($R^2=0.401$ versus $0.354$). Finally, a hybrid architecture co-training EnsembleEGNN with the BERT model achieves the highest overall accuracy across both random ($R^2=0.538$) and structural holdout evaluations ($R^2=0.444$). These results demonstrate that encoding conformational ensembles into latent representations improves predictions for properties governed by thermodynamics.
CommentsAccepted to Graph Foundation Models workshop at ICML '26. Contains 18 pages, 4 figures, 3 tables, 2 SI items