arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33122q-bio.BM

PHL:用于蛋白质-蛋白质结合亲和力预测的持久超有向图学习

PHL: Persistent Hyperdigraph Learning for Protein-Protein Binding Affinity Prediction

发表机构佛罗里达大学 · 阿肯色大学
查看机构详情
  • University of Florida(佛罗里达大学)
  • University of Arkansas(阿肯色大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingjian Xu, Chunmei Wang, Jiahui Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出持久超有向图学习(PHL)及其无矩阵变体MFPHL,直接建模方向性多体相互作用,通过随机迹估计大幅降低计算成本,在蛋白质结合亲和力预测中实现约百倍加速,精度与谱方法相当。

中文摘要 AI 辅助

持久同调与持久拉普拉斯算子已成为大分子结合相互作用的有效描述符,后者编码了超越调和子空间的多尺度几何与谱信息。然而,两者均基于无向图与成对接触,而决定结合的相互作用往往是方向性的,且同时涉及两个以上原子。我们引入了持久超有向图学习(PHL),直接表示这些有向的多体相互作用。由于超有向图拉普拉斯谱在更高拓扑阶上计算成本高昂,我们进一步利用随机迹估计,开发了一种无矩阵公式(MFPHL),通过稀疏边界算子求值的基于探针的二次型来估计拉普拉斯迹统计量,无需组装或对角化。达到目标相对迹估计精度所需的探针数量与矩阵维度无关,且每探针成本随算子稀疏度缩放。在蛋白质-蛋白质结合亲和力基准上,MFPHL相对于基于特征值的流程,将特征生成时间减少了约两个数量级。其预测精度在训练种子变异范围内与P2P野生型集合上的谱PHL相当,但在两个较大数据集上较低。这种依赖于数据集的权衡使高阶超有向图特征在实际应用中变得可行。

英文摘要

Persistent homology and persistent Laplacians have become effective descriptors of large molecular binding interactions, the latter encoding multiscale geometry and spectral information beyond the harmonic subspace. Both, however, are built on undirected graphs and pairwise contracts, while the interactions that determine binding are often directional and involve more than two atoms at once. We introduce persistent hyperdigraph learning (PHL), which represents these directed, many-body interactions directly. Because hyperdigraph Laplacian spectra become costly to compute at higher topological orders, we further use stochastic trace estimation to develop a matrix-free formulation(MFPHL) that estimates Laplacian trace statistics from probe-based quadratic forms evaluated through sparse boundary operators, without assembling or diagonalizing. The number of probes required for a target relative accuracy of the trace estimate is independent of matrix dimension, and per-probe cost scales with operator sparsity. On protein-protein binding affinity benchmarks, MFPHL reduces feature-generation time by roughly two orders of magnitude relative to the eigenvalue-based pipeline. Its predictive accuracy matches that of spectral PHL on the P2P wild-type set within training-seed variability but is lower on the two larger datasets. This dataset-dependent trade-off brings higher-order hyperdigraph features within practical reach.

↑