arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过强化稀疏自编码器特征的内在无序蛋白区域生成建模

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff

arXiv 2610.02189首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出IDiom模型和RL-SAE方法,用于内在无序蛋白区域的可解释设计,显著提升功能特征激活率。

AI 中文摘要

内在无序蛋白区域(IDRs)在转录调控、信号转导和亚细胞定位等细胞过程中发挥核心作用,然而其功能设计仍然具有挑战性。基于结构的设计方法不直接适用于IDRs,现有的蛋白质语言模型是在全长蛋白质序列上训练的,因此学习到的先验偏向于折叠结构域。在此,我们提出了IDiom,一种自回归蛋白质语言模型,在IDiom-DB上训练,该数据集包含从AlphaFold数据库整理的5400万个预测IDRs。IDiom生成多样化的序列,能够重现天然IDRs的组成、模式、基序和预测无序性。为了控制与功能相关的序列模式,我们还引入了带稀疏自编码器特征的强化学习(RL-SAE),这是一种后训练方法,奖励生成激活特定特征集的序列。在八个IDR设计任务中,RL-SAE序列平均激活了30个目标特征中的90%,而激活引导仅为24%。我们证明,与引导和监督微调相比,RL-SAE提高了生成IDRs的预测亚细胞定位和转录活性,并能够在单个序列中组合与不同生物学功能相关的特征。因此,IDiom和RL-SAE通过显式控制功能相关的序列特征,实现了可解释且可组合的IDR设计。更广泛地说,RL-SAE可以扩展到其他蛋白质设计场景,其中可解释特征提供了有用的设计目标。代码可在https://this URL获取。

英文摘要

Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑