arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03978cs.LGq-bio.BM

学习潜在蛋白质语言用于自回归生成

Learning Latent Protein Languages for Autoregressive Generation

Mahdi Pourmirzaei, Farzaneh Esmaili, Amir Ziashahabi, Mohammadreza Pourmirzaei, Dong Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出两种潜在蛋白质语言(PLL和SLL),通过自回归Transformer预训练,在序列生成、结构预测和骨架生成上显著优于氨基酸模型,并支持快速采样。

中文摘要 AI 辅助

自回归Transformer在蛋白质序列和结构生成方面仍然相对较弱。我们研究了目标表示的作用:氨基酸令牌编码残基身份而不具备明确的上下文语义,而骨架坐标在我们的框架中需要离散表示。我们引入了两种学习到的潜在蛋白质语言。蛋白质潜在语言(PLL)基于冻结的ESM-2编码器将序列映射到包含4,096个状态的上下文字母表,每个残基对应一个令牌。结构潜在语言(SLL)在保留解码到骨架坐标的同时,通过辅助序列和置信度监督来适配GCP-VQVAE Lite。我们分别使用下一个令牌预测在PLL和SLL令牌上预训练自回归Transformer模型,得到PLLM和SLLM。在匹配的下游序列训练下,PLLM的拟合计算缩放指数为0.038,而氨基酸自回归模型为0.020。在无条件序列生成中,PLLM将低于启发式1.5比特残基组成熵阈值的样本比例相对于氨基酸模型在不同采样温度下减少了54%。对于序列到结构预测,在匹配训练下,用SLL替换原始GCP-VQVAE Lite分词器使最佳验证困惑度降低了34%。对于长蛋白质,在我们的测量中,潜在令牌采样比基于MSA的AlphaFold2快约1,000倍。在骨架生成中,SLLM在多样性和新颖性方面与其他生成模型相比表现良好。我们还观察到早期迹象表明,在推理时采样中使用SLLM的内部令牌置信度可以提高序列到结构预测的质量,超越单个解码样本。这些结果将学习到的潜在蛋白质语言定位为自回归Transformer扩展和蛋白质生成中推理时采样的有前景的基础。

英文摘要

Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework. We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, with one token per residue. Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while retaining decoding to backbone coordinates. We separately pretrain autoregressive transformer models on PLL and SLL tokens using next-token prediction, yielding PLLM and SLLM. Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model. In unconditional sequence generation, PLLM reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across sampling temperatures. For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training. For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in our measurements. In backbone generation, SLLM compares favorably with other generative models on diversity and novelty. We also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample. These results position learned latent protein languages as a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.

发表机构

  • University of Missouri(密苏里大学)
  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑