arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04850cs.LGcs.AI

PIT-GCL:基于拓扑图对比学习的蛋白质相互作用

PIT-GCL: Protein Interaction using Topological Graph Contrastive Learning

Jae Won Choi, Ryoonki Hong, Alan Liang, Manjula Adiveppa Wader, Bingsong Zeng, Peiyang Tang, Longwei Liu, Ruishan Liu

首次发表
浏览论文内容

中文总结 AI 辅助

PIT-GCL提出双塔结构感知框架,结合序列、Cα点云和持续同调特征,通过交叉注意力软对接与对比学习,在多个蛋白质结合预测基准上取得领先,并支持大规模筛选。

中文摘要 AI 辅助

蛋白质结合预测是靶点识别、治疗性结合物设计和大规模筛选的核心任务,但由于结合依赖于序列、三维几何结构以及整体结构组织,该任务仍具挑战性。近期如AlphaFold3和Boltz-2等折叠模型显著提升了结构预测性能,但其置信度输出(pLDDT、pTM、ipTM)并非专门针对二元结合预测设计,且专用的结构感知预测器通常需要结合态复合物结构,而这类结构在筛选规模下难以获得。我们提出PIT-GCL,一种双塔结构感知框架,分别从氨基酸序列、Cα点云和全局持续同调描述符独立编码每个蛋白质。每个塔将残基ESM-2嵌入与由Vietoris-Rips过滤的H0和H1持续景观计算得到的拓扑摘要相结合,并通过结构感知Transformer处理所得令牌,其中成对Cα距离作为学习到的注意力偏置输入。随后,双向交叉注意力模块在两个蛋白质表示之间执行潜在空间软对接,模型通过二元交叉熵与NT-Xent对比损失的联合目标进行训练。在三个二元交互预测基准上——PPIRef上的通用PPI、STAG上的TCR-pMHC结合以及PPB-Affinity上的全链对——PIT-GCL在我们的评估中于通用PPI上优于代表性的基于序列、结构感知和任务特定的基线,并且是唯一在PPB-Affinity上高于随机水平的方法;在TCR-pMHC上,它在固定决策阈值下领先,但被一个任务特定的序列模型超越。由于每个蛋白质在第一阶段被独立编码,其表示可预先计算并在候选对之间重复使用,这便于大规模筛选。

英文摘要

Protein binding prediction is central to target identification, therapeutic binder design, and large scale screening, yet remains challenging because binding depends on sequence, three dimensional geometry, and global structural organization. Recent folding models such as AlphaFold3 and Boltz-2 have substantially improved structure prediction, but their confidence outputs (pLDDT, pTM, ipTM) are not specifically designed for binary binding prediction, and dedicated structure aware predictors often require bound complex structures that are unavailable at screening scale. We introduce PIT-GCL, a dual tower structure aware framework that encodes each protein independently from its amino acid sequence, Cα point cloud, and a global persistent homology descriptor. Each tower combines residue ESM-2 embeddings with a topological summary computed from the H0 and H1 persistence landscapes of a Vietoris-Rips filtration, and processes the resulting tokens with a structure aware Transformer in which pairwise Cα distances enter as a learned attention bias. A bidirectional cross attention module then performs latent space soft docking between the two per-protein representations, and the model is trained with a combined binary cross entropy and NT-Xent contrastive objective. On three binary interaction prediction benchmarks, general PPI on PPIRef, TCRpMHC binding on STAG, and whole chain pairs on PPB-Affinity, PIT-GCL outperforms representative sequence based, structure aware, and task specific baselines on general PPI under our evaluation, and is the only method above chance on PPB-Affinity; on TCR-pMHC it leads at a fixed decision threshold but is outranked by a task specific sequence model. Because each protein is encoded independently in the first phase, its representation can be precomputed and reused across candidate pairs, which is convenient for large scale screening.

发表机构

  • University of Southern California(南加州大学)
  • The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑