arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.10244cs.AI

PepTriX:一种通过蛋白质语言模型实现可解释肽分析的框架

PepTriX: A Framework for Explainable Peptide Analysis through Protein Language Models

  • Center for Artificial Intelligence in Public Health Research (ZKI-PH), Robert Koch Institute(人工智能与公共健康研究所以及罗伯特·科赫研究所)
  • Department of Mathematics and Computer Science, Free University of Berlin(数学与计算机科学系,柏林自由大学)

机构由 AI 辅助整理,请以论文原文为准。

Vincent Schilling, Akshat Dubey, Georges Hattab

更新

AI总结:

针对肽分类任务中传统方法通用性受限、蛋白质语言模型微调成本高且可解释性差的问题,提出PepTriX框架,整合1D序列嵌入与3D结构特征,实现跨任务优异预测并提供可解释生物学见解。

AI中文摘要:

肽分类任务(如预测毒性和HIV抑制作用)是生物信息学和药物发现的基础。传统方法严重依赖手工编码的一维(1D)肽序列,这可能限制跨任务和数据集的通用性。近年来,蛋白质语言模型(PLMs)(如ESM-2和ESMFold)已展现出强大的预测性能,但面临两个关键挑战:一是微调计算成本高,二是复杂的潜在表示阻碍领域专家的可解释性。此外,许多框架针对特定肽分类类型开发,缺乏通用性,限制了将模型预测与生物学相关基序和结构特性联系起来的能力。为此,我们提出PepTriX框架,通过对比训练和跨模态共注意力增强的图注意力网络,整合一维(1D)序列嵌入和三维(3D)结构特征。PepTriX可自动适应不同数据集,生成任务特异性肽向量并保留生物学合理性。经领域专家评估,PepTriX在多个肽分类任务中表现优异,能提供驱动预测的结构和生物物理基序的可解释见解,从而弥合性能驱动的肽级模型(PLMs)与肽研究领域理解之间的差距。

英文摘要:

Peptide classification tasks, such as predicting toxicity and HIV inhibition, are fundamental to bioinformatics and drug discovery. Traditional approaches rely heavily on handcrafted encodings of one-dimensional (1D) peptide sequences, which can limit generalizability across tasks and datasets. Recently, protein language models (PLMs), such as ESM-2 and ESMFold, have demonstrated strong predictive performance. However, they face two critical challenges. First, fine-tuning is computationally costly. Second, their complex latent representations hinder interpretability for domain experts. Additionally, many frameworks have been developed for specific types of peptide classification, lacking generalization. These limitations restrict the ability to connect model predictions to biologically relevant motifs and structural properties. To address these limitations, we present PepTriX, a novel framework that integrates one dimensional (1D) sequence embeddings and three-dimensional (3D) structural features via a graph attention network enhanced with contrastive training and cross-modal co-attention. PepTriX automatically adapts to diverse datasets, producing task-specific peptide vectors while retaining biological plausibility. After evaluation by domain experts, we found that PepTriX performs remarkably well across multiple peptide classification tasks and provides interpretable insights into the structural and biophysical motifs that drive predictions. Thus, PepTriX offers both predictive robustness and interpretable validation, bridging the gap between performance-driven peptide-level models (PLMs) and domain-level understanding in peptide research.

↑