arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

蛋白质语言模型中特定任务与数据集的信息

Task- and dataset-specific information in protein language models

Roman Joeres, Ilya Senatorov, Anastasia Kolchina, Dietrich Klakow, Olga V. Kalinina

arXiv 2608.12090首次发表:更新:

发表机构

Helmholtz Institute for Pharmaceutical Research Saarland; Saarland University; Saarland Informatics Campus; Pharma Science Hub; Medical Faculty, University Hospital Saarland(萨尔兰亥姆霍兹药物研究所; 萨尔兰大学; 萨尔兰信息学园区; 制药科学中心; 萨尔兰大学医院医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究分析13个蛋白质语言模型在15个下游任务上的表现,发现其最后一层嵌入并非最优,任务与数据集特性影响各层嵌入的信息有效性,人工蛋白质任务会显著降低模型性能。

AI 中文摘要

蛋白质语言模型(Protein Language Models, PLMs)已将自然语言处理领域的最新进展迁移至计算生物学,这类在大量蛋白质序列语料库上训练得到的模型,被广泛用于将氨基酸序列转换为可应用于各类下游任务(Downstream Tasks, DTs)的隐空间嵌入。根据普遍共识,人们通常使用模型最后一层的嵌入,但对模型的内部行为仍知之甚少。我们分析了13个PLMs在来自11个数据集的15个DTs上的表现,以探究PLMs中间层生成的嵌入的信息含量。我们在各层生成的嵌入上训练探测模型,对比其性能,并计算它们所跨越的隐空间的特征,以估计其中包含的信息,结果发现PLMs的最后一层极少包含能在下游任务中取得最佳结果的嵌入。此外,我们确定了DTs与PLMs各层中用于预测该任务的相关信息分布之间存在关联,例如,预训练目标与预测单个残基属性的目标之间的相似性,会使得PLMs各层对这类任务的理解逐步提升;而对于全蛋白质任务,我们观察到决定PLMs在DT上表现的是数据集而非任务本身,包含深度突变扫描(Deep Mutational Scan, DMS)数据的数据集,使用PLMs浅层的嵌入表现更好,包含多样天然蛋白质的数据集则会发现PLMs深层的嵌入最有用。另外,我们还发现,当引入针对人工蛋白质的任务时,PLMs的性能会显著下降。

英文摘要

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs). By consensus, embeddings from the models' last layers are used, while the models' internal behavior remains poorly understood. We analyzed 13 PLMs across 15 DTs and 9 datasets to assess the value of embeddings from intermediate PLM layers. We trained probe models on embeddings from each layer, compared their performance, and showed that the last layers of PLMs rarely produced embeddings that led to the best results on downstream tasks. Furthermore, we identified a connection between how models learn a certain DT and the similarity between that DT and the pre-training objective. For example, for residue-level downstream tasks, we observed a steady increase in performance across almost all PLM layers, which we attributed to their similarity to most PLMs' pre-training objectives. To allow the community to capitalize on our findings, we provide PLMSommelier, a Python package that automatically identifies the best PLM layer for a given DT with ~98% accuracy and creates a truncated model using only the early layers up to the best-performing layer. This will help users save time and memory during inference and yield better predictive performance.

Comments36 pages, 14 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑