预测器引导的潜空间密码子优化以最大化蛋白质表达
Predictor-Guided Latent Space Codon Optimization for Maximizing Protein Expression
- Johnson & Johnson Innovative Medicine(强生创新医药)
- Ellison Institute of Technology Oxford(埃利森技术研究所牛津分校)
- New Theory AI(新理论人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对密码子优化难题,提出潜空间密码子优化(LSCO),利用预训练mRNA语言模型将离散问题连续化,结合预测器目标、自由能正则化、自然性先验和约束解码,在湿实验抗体数据上超越现有基线。
AI中文摘要:
密码子优化,即选择同义密码子以提高mRNA翻译效率和蛋白质表达的过程,对于治疗性蛋白质生产和mRNA疫苗至关重要,但它仍然是一个难题。设计空间是离散且组合巨大的,排除了基于梯度的方法,现有工具依赖于启发式代理(如密码子适应指数或GC含量),这些代理不能很好地捕捉真实表达。我们引入了潜空间密码子优化(LSCO),通过将序列映射到预训练的mRNA语言模型的潜空间中,将这一离散问题重新表述为连续问题,从而实现高效的基于梯度的搜索。LSCO结合了四个组成部分:来自不确定性感知预测器的数据驱动表达目标、用于结构稳定性的最小自由能正则化器、来自蛋白质到密码子反向翻译模型的自然性先验,以及用于蛋白质保真度的约束解码。在一个真实世界的湿实验室抗体表达数据集上,LSCO在预测表达方面优于简单的基于频率的以及现代深度生成基线,同时保留了合适的生物物理特性。
英文摘要:
Codon optimization, the process of selecting synonymous codons to improve mRNA translation efficiency and protein expression, is central to therapeutic protein production and mRNA vaccines, yet it remains a hard problem. The design space is discrete and combinatorially large, precluding gradient-based methods, and existing tools rely on heuristic proxies (e.g., Codon Adaptation Index or GC-content) that poorly capture true expression. We introduce Latent-Space Codon Optimization (LSCO), which recasts this discrete problem as a continuous one by mapping sequences into the latent space of a pretrained mRNA language model, enabling efficient gradient-based search. LSCO combines four components: a data-driven expression objective from an uncertainty-aware predictor, a Minimum-Free-Energy regularizer for structural stability, a naturalness prior from a protein-to-codon back-translation model, and constrained decoding for protein fidelity. On a real-world, wet-lab antibody expression dataset, LSCO outperforms simple frequency-based, as well as modern deep generative baselines in predicted expression, while retaining suitable biophysical properties.