arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25334cs.LGcs.CV

MT-ProtBERT:面向稀缺数据下固有无序蛋白质分类的多任务学习ProtBERT

MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data

Jian Sun, Kingshuk Ghosh, Lilianna Houston, Mohammad H. Mahoor

首次发表
浏览论文内容

中文总结 AI 辅助

针对固有无序蛋白质数据稀缺导致分类困难的问题,提出多任务ProtBERT(MT-ProtBERT),集成动态窗口掩码、多尺度卷积和辅助任务,在磷酸化位点预测和蛋白质压缩预测上优于PARROT。

中文摘要 AI 辅助

固有无序蛋白质(IDPs)与折叠蛋白质不同,它们具有动态性,缺乏稳定的三维构象,并且相似蛋白质之间的序列相似性较低。IDPs的构象异质性——虽然对其多样功能有益——限制了使用传统实验工具来确定其构象。实验难度加上低序列相似性导致数据稀缺,使得对相似或不相似的IDPs进行分类/检测变得困难,这一任务与理解生物学和进化相关。我们使用多任务ProtBERT(MT-ProtBERT)应对这一挑战,这是ProtBERT的多任务扩展,专为低数据场景设计。MT-ProtBERT集成了动态窗口掩码、多尺度一维卷积分类器(MS-Conv1D)以及辅助目标,这些目标联合优化掩码语言建模和基于生物化学的任务。我们在有限数据下评估了该框架在两个任务上的表现:(i)短序列和小数据集中的磷酸化位点预测(S/T/Y),以及(ii)在两个小数据集(684和530个序列)上的蛋白质压缩预测,包括长度与典型无序区域相当的序列。MT-ProtBERT在所有任务上始终优于基于RNN的IDP专用模型PARROT。这些结果表明,结合自监督和基于生物化学的任务以及多尺度学习,能够在数据稀缺条件下实现对无序蛋白质的稳健建模。

英文摘要

Intrinsically disordered proteins (IDPs) differ from folded proteins in that they are dynamic, lack a stable three-dimensional conformation, and have low sequence similarity between similar proteins. The conformational heterogeneity of IDPs - while beneficial for their diverse functions - limits the use of traditional experimental tools to determine their conformation. The experimental difficulty, along with low sequence similarity, results in data scarcity, and makes it difficult to classify/detect IDPs that are similar or dissimilar, a task relevant to understand biology and evolution. We address this challenge using Multi-task ProtBERT (MT-ProtBERT), a multi-task extension of ProtBERT tailored for low-data regimes. MT-ProtBERT integrates Dynamic Window Masking, a Multi-Scale 1D Convolutional classifier (MS-Conv1D), and auxiliary objectives that jointly optimize masked language modeling and biochemistry-informed tasks. We evaluate this framework on two tasks under limited data: (i) phosphorylation site prediction (S/T/Y) in short sequences and small datasets, and (ii) protein compaction prediction on two small datasets (684 and 530 sequences), including sequences comparable in length to typical disordered regions. MT-ProtBERT consistently outperforms PARROT, an RNN-based IDP-specific model, across all tasks. These results demonstrate that combining self-supervised and biochemistry-informed tasks, and multi-scale learning enables robust modeling of unstructured proteins under data scarcity.

发表机构

  • University of Denver(丹佛大学)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑