发表机构
ShanghaiTech University(上海科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ProtLingo框架,通过条件局部记忆与稀疏专家路由机制构建高效蛋白质语言模型,在1.5亿参数规模下于多项蛋白质预测任务取得竞争力性能,兼具参数效率与长程结构表征能力。
AI 中文摘要
蛋白质执行多种细胞功能,即便单个氨基酸替换也能改变其稳定性、活性或分子间相互作用。蛋白质语言模型(PLMs)为从未标记序列中建模这类序列-功能关系提供了可扩展方法,但增大稠密Transformer主干网络的规模往往会带来巨大计算成本,且无法持续提升突变敏感型预测性能。本文提出ProtLingo,这是一种高效PLM框架,它在预训练单序列主干网络基础上增加了条件局部记忆与稀疏专家路由机制。ProtLingo将上下文残基表征映射为路由特定的离散编码,将居中局部窗口组合为潜在N元语法地址,并检索与重复出现的局部序列上下文相关的可复用残差信号;同时,部分选定的前馈块被升级为带有共享专家与路由专家的稀疏混合专家层,在仅激活部分参数的情况下实现依赖残基的计算。在蛋白质适应性预测、FLIP基准测试及监督式接触预测任务上的实验表明,采用1.5亿参数规模主干网络的ProtLingo实现了极具竞争力的性能,包括在突变效应预测上表现出优异的参数效率,且保留了长程结构表征能力。
英文摘要
Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transformer backbones often brings substantial computational cost without consistently improving mutation-sensitive prediction. We introduce ProtLingo, an efficient PLM framework that augments a pretrained single-sequence backbone with conditional local memory and sparse expert routing. ProtLingo maps contextual residue representations into route-specific discrete codes, composes centered local windows into latent $N$-gram addresses, and retrieves reusable residual signals associated with recurring local sequence contexts. In parallel, selected feed-forward blocks are upcycled into sparse Mixture-of-Experts layers with shared and routed experts, enabling residue-dependent computation while activating only a subset of parameters. Experiments on protein fitness prediction, FLIP benchmarks, and supervised contact prediction show that ProtLingo achieves competitive performance with a 150M-scale backbone, including strong parameter efficiency on mutation-effect prediction and preserved long-range structural representations.