发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出IDiom模型和RL-SAE方法,用于内在无序蛋白区域的可解释设计,显著提升功能特征激活率。
AI 中文摘要
内在无序蛋白区域(IDRs)在转录调控、信号转导和亚细胞定位等细胞过程中发挥核心作用,然而其功能设计仍然具有挑战性。基于结构的设计方法不直接适用于IDRs,现有的蛋白质语言模型是在全长蛋白质序列上训练的,因此学习到的先验偏向于折叠结构域。在此,我们提出了IDiom,一种自回归蛋白质语言模型,在IDiom-DB上训练,该数据集包含从AlphaFold数据库整理的5400万个预测IDRs。IDiom生成多样化的序列,能够重现天然IDRs的组成、模式、基序和预测无序性。为了控制与功能相关的序列模式,我们还引入了带稀疏自编码器特征的强化学习(RL-SAE),这是一种后训练方法,奖励生成激活特定特征集的序列。在八个IDR设计任务中,RL-SAE序列平均激活了30个目标特征中的90%,而激活引导仅为24%。我们证明,与引导和监督微调相比,RL-SAE提高了生成IDRs的预测亚细胞定位和转录活性,并能够在单个序列中组合与不同生物学功能相关的特征。因此,IDiom和RL-SAE通过显式控制功能相关的序列特征,实现了可解释且可组合的IDR设计。更广泛地说,RL-SAE可以扩展到其他蛋白质设计场景,其中可解释特征提供了有用的设计目标。代码可在https://this URL获取。
英文摘要
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.