结合大规模平行筛选优化RNA产量的深度学习方法
Optimizing RNA yield using deep neural networks coupled to massively parallel screening
AI总结:
本研究结合大规模平行测序与卷积神经网络,构建了可高分辨率预测RNA产量的深度学习框架,其预测与实验产量的皮尔逊相关系数达0.94,可用于RNA序列的生产前优先排序,降低成本并加速RNA工程。
AI中文摘要:
基于信使RNA(mRNA)的疗法已成为疫苗、蛋白替代疗法和癌症免疫治疗的强大平台。mRNA开发的关键瓶颈是经济地大量生产RNA,这通过体外转录(IVT)反应产生的RNA产量来衡量。然而,启动子相邻的DNA序列如何影响RNA产量的特征仍未被充分阐明。在此,我们提出一种集成深度学习框架,利用大规模平行下一代测序(NGS)分析来测量大型序列空间中的RNA产量。设计了包含10^5个随机寡核苷酸序列的文库,以在确定的结构背景下系统探索序列多样性。使用Illumina测序平行量化DNA和RNA丰度,实现大规模序列-产量关系的高分辨率测量。将序列进行独热编码,用于训练采用卷积神经网络架构的深度学习模型。该模型在留出测试集上预测与实验测量的RNA产量的皮尔逊相关系数达到0.94,证明其在不同序列背景下具有强泛化能力。重要的是,训练后的模型可部署到生产环境中,通过预测的IVT产量对新型RNA序列设计进行评分和排序,实现对最具可制造性候选物的成本效益高的实验前优先排序。该框架建立了一种可扩展的数据驱动型DNA和RNA序列优化方法,广泛适用于疫苗抗原设计、治疗蛋白递送和合成生物学。通过将高通量实验与先进的深度学习建模相结合,它显著降低了筛选成本并缩短了RNA工程周期时间。
英文摘要:
Messenger RNA (mRNA)-based therapeutics have emerged as a powerful platform for vaccines, protein replacement therapies, and cancer immunotherapy. A critical bottleneck in mRNA development is manufacturing large quantities of RNA economically, as measured by RNA yield emerging from an in vitro transcription (IVT) reaction. However, how promoter-adjacent DNA sequences influence RNA yield remains poorly characterized. Here, we present an integrated deep learning framework that leverages massively parallel next-generation sequencing (NGS) assays to measure RNA yield across large sequence spaces. A library of 10^5 randomized oligonucleotide sequences was designed to systematically explore sequence diversity within a defined structural context. DNA and RNA abundances were quantified in parallel using Illumina sequencing, enabling high-resolution measurement of sequence-to-yield relationships at scale. Sequences were one-hot encoded and used to train deep learning models, using a convolutional neural network architecture. The model achieved a Pearson correlation of 0.94 between predicted and experimentally measured RNA yield on a held-out test set, demonstrating strong generalization across diverse sequence contexts. Importantly, the trained model can be deployed in a production environment to score and rank novel RNA sequence designs by predicted IVT yield, enabling cost-effective, pre-experimental prioritization of the most manufacturable candidates. This framework establishes a scalable, data-driven approach to DNA and RNA sequence optimization, with broad applicability to vaccine antigen design, therapeutic protein delivery, and synthetic biology. By integrating high-throughput experimentation with advanced deep learning modeling, it significantly reduces screening costs and accelerates RNA engineering cycle times.