arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于表征化学和结构贡献对水溶性的加性MLP-GNN框架

An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility

Sampreeti Bhattacharya, Arkaprava Roy

arXiv 2607.02212首次发表:更新:

AI 中文总结

提出加性深度学习框架,分别用MLP编码理化描述符、GNN编码分子图拓扑,通过加性模型组合预测水溶性,实现化学与结构贡献的分离解释,并在AqSolDB和BigSolDB2数据集上取得竞争性能。

AI 中文摘要

水溶性是早期药物发现中的关键性质,但大多数预测模型将物理化学描述符和分子图信息合并为单一表示,模糊了预测是由全局化学、分子结构还是两者共同驱动。我们提出一个加性深度学习框架,在训练过程中保持这两类信息分离:物理化学描述符由多层感知机(化学分支)编码,分子图拓扑由图神经网络(结构分支)编码,两个输出仅在预测阶段通过加性模型(可选乘性交互)组合。这种设计提供了化学和结构组分的直接分解,可在训练后单独检查。此外,在较大的AqSolDB数据集上预训练并在较小的BigSolDB2数据集上微调,显著提高了准确性并减少了运行间变异,表明从数据丰富场景中学到的特征具有泛化性。我们进一步使用分支输出的最佳线性投影、跨溶解度类别的分子级嵌入摘要以及聚集在官能团上的原子级GNNExplainer掩码来解释拟合模型。这些分析表明,化学分支与熟悉的物理化学描述符对齐,而结构分支捕捉与溶解度相关的图拓扑和官能团模式。在两个数据集上,该框架在保持竞争性预测性能的同时,使化学和结构信息的独特作用更加透明。

英文摘要

Aqueous solubility is a key property in early-stage drug discovery, but most predictive models merge physicochemical descriptors and molecular graph information into a single representation, obscuring whether a prediction is driven by global chemistry, molecular structure, or both. We present an additive deep-learning framework that keeps these two sources of information separate throughout training: physicochemical descriptors are encoded by a multilayer perceptron (the chemical branch) and molecular graph topology by a graph neural network (the structural branch), with the two outputs combined only at the prediction stage through an additive model with an optional multiplicative interaction. This design provides a direct decomposition of chemical and structural components that can be examined separately after training. Furthermore, pretraining on the larger AqSolDB dataset and fine-tuning on the smaller BigSolDB2 dataset substantially improve accuracy and reduce run-to-run variations, indicating generalizability of the learned features from the data-rich settings. We further interpret the fitted model using best linear projections of the branch outputs, molecule-level embedding summaries across solubility classes, and atom-level GNNExplainer masks aggregated over functional groups. These analyses show that the chemical branch aligns with familiar physicochemical descriptors, while the structural branch captures graph-topological and functional-group patterns associated with solubility. Across both datasets, the framework attains competitive predictive performance while making the distinct roles of chemical and structural information more transparent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑