arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03098cs.AIcs.LG

预测器引导的潜空间密码子优化以最大化蛋白质表达

Predictor-Guided Latent Space Codon Optimization for Maximizing Protein Expression

  • Johnson & Johnson Innovative Medicine(强生创新医药)
  • Ellison Institute of Technology Oxford(埃利森技术研究所牛津分校)
  • New Theory AI(新理论人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

Alberto Caron, Tianyu Cui, Dmytro S. Lituiev, Mangal Prakash, Artem Moskalev, Amina Mollaysa, Bo Zhai, Hirsh Nanda, Daniel M. Poole, Zhongyin Liu, Iman Farasat,… 展开作者

Alberto Caron, Tianyu Cui, Dmytro S. Lituiev, Mangal Prakash, Artem Moskalev, Amina Mollaysa, Bo Zhai, Hirsh Nanda, Daniel M. Poole, Zhongyin Liu, Iman Farasat, Robert Davidson, Nikolay V. Manyakov, Tommaso Mansi, Scott Oloff, Rui Liao

AI总结:

针对密码子优化难题,提出潜空间密码子优化(LSCO),利用预训练mRNA语言模型将离散问题连续化,结合预测器目标、自由能正则化、自然性先验和约束解码,在湿实验抗体数据上超越现有基线。

AI中文摘要:

密码子优化,即选择同义密码子以提高mRNA翻译效率和蛋白质表达的过程,对于治疗性蛋白质生产和mRNA疫苗至关重要,但它仍然是一个难题。设计空间是离散且组合巨大的,排除了基于梯度的方法,现有工具依赖于启发式代理(如密码子适应指数或GC含量),这些代理不能很好地捕捉真实表达。我们引入了潜空间密码子优化(LSCO),通过将序列映射到预训练的mRNA语言模型的潜空间中,将这一离散问题重新表述为连续问题,从而实现高效的基于梯度的搜索。LSCO结合了四个组成部分:来自不确定性感知预测器的数据驱动表达目标、用于结构稳定性的最小自由能正则化器、来自蛋白质到密码子反向翻译模型的自然性先验,以及用于蛋白质保真度的约束解码。在一个真实世界的湿实验室抗体表达数据集上,LSCO在预测表达方面优于简单的基于频率的以及现代深度生成基线,同时保留了合适的生物物理特性。

英文摘要:

Codon optimization, the process of selecting synonymous codons to improve mRNA translation efficiency and protein expression, is central to therapeutic protein production and mRNA vaccines, yet it remains a hard problem. The design space is discrete and combinatorially large, precluding gradient-based methods, and existing tools rely on heuristic proxies (e.g., Codon Adaptation Index or GC-content) that poorly capture true expression. We introduce Latent-Space Codon Optimization (LSCO), which recasts this discrete problem as a continuous one by mapping sequences into the latent space of a pretrained mRNA language model, enabling efficient gradient-based search. LSCO combines four components: a data-driven expression objective from an uncertainty-aware predictor, a Minimum-Free-Energy regularizer for structural stability, a naturalness prior from a protein-to-codon back-translation model, and constrained decoding for protein fidelity. On a real-world, wet-lab antibody expression dataset, LSCO outperforms simple frequency-based, as well as modern deep generative baselines in predicted expression, while retaining suitable biophysical properties.

↑