发表机构
Nanyang Technological University; Sun Yat-sen University; Zhejiang University; Carnegie Mellon University; Shanghai Institute of Optics and Fine Mechanics; Zhongnan Hospital, Wuhan University; University of Southern California(南洋理工大学; 中山大学; 浙江大学; 卡内基梅隆大学; 上海光学精密机械研究所; 武汉大学中南医院; 南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出线性适应度子空间假设,并开发子空间引导进化搜索(SGES),在少量标记样本下高效估计适应度方向,显著提升蛋白质定向进化的预测与搜索效率。
AI 中文摘要
模型引导的定向进化旨在有限的oracle预算下识别高适应度的蛋白质变体。蛋白质语言模型(PLMs)为此任务提供了丰富的表示,但任务无关的零样本得分可能与目标测定不对齐,而在高维嵌入空间中的监督搜索可能使代理建模和不确定性估计样本效率低下。我们提出线性适应度子空间(LFS)假设:在突变诱导的残基级表示变化中,一组紧凑的、测定特异的方向使得适应度变化可以从少量标记变体中线性地获得。这是一个局部的、可通过监督恢复的陈述,而非声称蛋白质适应度景观或全局PLM几何结构普遍是线性的。基于这一观察,我们引入了子空间引导进化搜索(SGES),它从一个小初始样本估计LFS,并在学习到的子空间中进行代理建模、不确定性估计和采集。在10个核心ProteinGym测定、87个扩展静态验证测定和18个测定的预算搜索评估中,SGES在适应度预测和搜索效率上优于零样本PLM和最近的ML引导蛋白质优化基线。与PCA、随机投影、标签洗牌PLS、经典突变特征和采集消融的对照比较进一步隔离了适应度对齐的位点-增量坐标的益处。
英文摘要
Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026