arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过正交投影解释蛋白质语言模型嵌入以用于蛋白质适应性预测

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal, Henrike O. Heyne, Nadja Klein, Bernhard Y. Renard

arXiv 2608.25548首次发表:更新:

发表机构

Hasso Plattner Institute, University of Potsdam; Broad Institute of MIT and Harvard; Humboldt University of Berlin; Karlsruhe Institute of Technology(波茨坦大学哈索·普拉特纳研究所; 麻省理工学院及哈佛大学布罗德研究所; 柏林洪堡大学; 卡尔斯鲁厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出通过正交投影技术去除蛋白质语言模型嵌入中的可解释生化特征效应,证明其编码与生化特性相关的模式并量化其对蛋白质适应性预测的贡献,该方法可迁移至其他问题场景。

AI 中文摘要

近年来,蛋白质语言模型(Protein Language Models, PLMs)在生物医学领域的应用日益广泛。它们的嵌入提供了蛋白质序列的丰富数值表示,在包括蛋白质适应性预测在内的多个下游任务中达到了最先进的性能。然而,PLM嵌入无法直接解释,因此它们编码了哪些特征仍不清楚。为深入了解蛋白质的哪些生化特性驱动了预测,我们利用正交投影技术,从嵌入中去除已知表格特征的线性效应,并将其扩展到高阶和交互效应。通过这种方式,我们从PLM嵌入中去除了可解释的生化特征的效应。在一项消融研究中,我们发现,对于仅使用嵌入训练以预测蛋白质适应性的下游分类器,这种操作会导致其性能下降。在另一项评估中,我们发现这些生化特征解释了该分类器预测中相当大一部分的方差。因此,我们可以证明PLM嵌入编码了与生化特性相关的模式,并量化了它们对预测蛋白质适应性的贡献。这种计算高效的方法不限于此处考虑的特征或嵌入,可轻松迁移到蛋白质适应性预测之外的问题场景。

英文摘要

Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. However, PLM embeddings are not directly interpretable and, thereby, it remains unclear what features they encode. To gain insight into which biochemical properties of the protein are driving the prediction, we leverage an orthogonal projection technique that removes linear effects of known tabular features from embeddings and extend it to high-order and interaction effects. In this way, we remove the effects of interpretable biochemical features from PLM embeddings. In an ablation study, we show that this leads to a decrease in performance for a downstream classifier trained only on the embeddings to predict protein fitness. In an additional evaluation, we find that these biochemical features explain a substantial part of the variance in the predictions of this classifier. Hence, we can show that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness. This computationally efficient approach is not limited to the features or embeddings considered here and is readily transferable to problem settings beyond protein fitness prediction.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑