arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PSLL:用于蛋白质-配体结合亲和力预测的持久层丛拉普拉斯学习

PSLL: Persistent Sheaf Laplacian Learning for Protein-Ligand Binding Affinity Prediction

Mushal Zia, Benjamin Jones, Guo-Wei Wei

arXiv 2609.05475首次发表:更新:

发表机构

University of Georgia; Michigan State University(佐治亚大学; 密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出持久层丛拉普拉斯学习(PSLL)框架,利用多尺度拓扑和原子电荷表示预测蛋白质-配体结合亲和力,在三个PDBbind基准上验证了强预测性能。

AI 中文摘要

准确的蛋白质-配体结合亲和力预测仍是计算药物发现中的核心挑战,这源于分子几何、物理化学相互作用和原子特异性电荷信息之间的复杂相互作用。在本工作中,我们提出了一种用于蛋白质-配体结合亲和力预测的持久层丛拉普拉斯学习(PSLL)框架。该方法通过将原子部分电荷纳入Vietoris-Rips和alpha复形过滤上的层丛限制映射,从三维蛋白质-配体复合物构建多尺度拓扑表示。为捕获化学多样化的蛋白质-配体相互作用,我们在PSLL框架中引入了元素特异性和类别特异性的原子对表示。从所得持久层丛拉普拉斯算子中提取的调和与非调和谱被用作分子描述符。为补充PSLL衍生的分子表示,我们结合了基于Transformer的蛋白质嵌入和SMILES衍生的配体描述符进行结合亲和力预测。所提出的多尺度PSLL模型的评分能力在三个广泛使用的PDBbind基准数据集(包括PDBbind-v2007、PDBbind-v2013和PDBbind-v2016)上,与现有最先进方法进行了验证。计算结果表明,所提出的PSLL模型在基准数据集上实现了强预测性能,突显了其作为可解释且数学基础扎实的框架的潜力,在分子机器学习和药物发现中具有良好的泛化能力。

英文摘要

Accurate prediction of protein-ligand binding affinity remains a central challenge in computational drug discovery due to the complex interplay among molecular geometry, physicochemical interactions, and atom-specific charge information. In this work, we introduce a Persistent Sheaf Laplacian learning (PSLL) framework for protein-ligand binding affinity prediction. The proposed approach constructs multiscale topological representations from three-dimensional protein-ligand complexes by incorporating atomic partial charges into sheaf restriction maps over Vietoris-Rips and alpha complex filtrations. To capture chemically diverse protein-ligand interactions, we introduce element-specific and category-specific atom-pair representations within the PSLL framework. Harmonic and non-harmonic spectra extracted from the resulting persistent sheaf Laplacians are used as molecular descriptors. To complement the PSLL-derived molecular representation, we incorporate transformer-based protein embeddings and SMILES-derived ligand descriptors for binding affinity prediction. The scoring power of the proposed multiscale PSLL model is validated against existing state-of-the-art methods on three widely used PDBbind benchmark datasets, including PDBbind-v2007, PDBbind-v2013, and PDBbind-v2016. The computational results indicate that the proposed PSLL model achieves strong predictive performance across benchmark datasets, highlighting its potential as an interpretable and mathematically grounded framework with promising generalizability for molecular machine learning and drug discovery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑