arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34654cs.AIq-bio.QM

蛋白质基础模型适应性预测的通用框架

A General Harness for Protein Foundation Model Fitness Prediction

Yang Tan, Qijia Tian, Gangyu Sun, Bozitao Zhong, Mingchen Li, Yuanxi Yu, Nanqing Dong, Liang Hong

首次发表
浏览论文内容

中文总结 AI 辅助

提出通用无训练检索增强框架VRH,融合模型得分与MSA证据及结构校正,在多个基准上显著提升蛋白质适应性预测,并构建了首个全面领先的VenusREM2模型。

中文摘要 AI 辅助

准确的适应性预测是蛋白质工程和理解序列-功能关系的核心。随着深度学习的进步,蛋白质基础模型(PFMs)已被广泛用于此任务。然而,最近的分析表明,这些模型共享反映其训练语料库的偏好,而不可靠的输入可能进一步扭曲适应性预测。家族特异性的进化证据和结构背景可以通过对模型得分提供互补约束来帮助解决这些限制,从而促使我们提出VenusREM-Harness(VRH),一个通用的、模型无关的、无需训练的检索增强突变框架。它根据模型不确定性将冻结的模型得分与多序列比对(MSA)证据融合,然后基于结构置信度和溶剂暴露应用门控背景校正和得分收缩。在来自ProteinGym、VenusMutHub和新整理的病毒基准VenusViroHub的1,211个测定和310万个测量变体中,所有71种配置在所有3个基准上的Spearman相关性平均提高了0.073,在5个指标上均有广泛提升。扩展分析将检索增益与模型-MSA偏好差异联系起来,评估了域级增益和免疫逃逸案例,并量化了计算加速。基于VRH构建的VenusREM2是首个在所有功能、分类群、MSA深度和突变深度类别中排名最高的模型,其在ProteinGym上的平均Spearman相关系数为0.556,比先前最佳结果高出0.038。

英文摘要

Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their training corpora, while unreliable inputs can further distort fitness predictions. Family-specific evolutionary evidence and structural context can help address these limitations by providing complementary constraints on model scores, motivating VenusREM-Harness (VRH), a general, model-agnostic, training-free Retrieval-Enhanced Mutation harness. It fuses frozen model scores with multiple sequence alignment (MSA) evidence according to model uncertainty, then applies gated background correction and score shrinkage based on structural confidence and solvent exposure. Across 1,211 assays and 3.1 million measured variants from ProteinGym, VenusMutHub, and the newly curated viral benchmark VenusViroHub, all 71 configurations improve Spearman correlation on all 3 benchmarks by 0.073 on average, with broad gains across 5 metrics. Extended analyses relate retrieval gains to model-MSA preference differences, assess domain-level gains and immune-escape cases, and quantify computational speedups. Built with VRH, VenusREM2 is the first to rank highest in all function, taxon, MSA-depth, and mutation-depth categories, with a ProteinGym Average Spearman of 0.556, 0.038 above the prior best.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑