条件感知学习实现跨多种化学修饰与检测条件下寡核苷酸熔解行为的稳健预测
Condition aware learning enables robust prediction of oligonucleotide melting behavior across diverse chemistries and assay conditions
- Cepheid(赛沛公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出条件感知核苷酸语言模型,结合序列表示与反应环境信息,实现跨化学修饰和实验条件的寡核苷酸熔解温度高精度预测,误差降低达25%。
AI中文摘要:
寡核苷酸熔解温度是核酸杂交的基本决定因素,支撑着分子诊断、聚合酶链式反应检测及众多其他生物技术应用的设计。然而,准确预测熔解行为仍然困难,因为这不仅取决于序列组成,还取决于现代检测设计中常用的实验条件和化学修饰。现有热力学模型依赖固定参数化,往往难以扩展到多样化的反应环境和核苷酸化学修饰。在此,我们展示了一种条件感知的核苷酸语言模型,能够在多种实验条件以及未修饰和化学修饰的寡核苷酸上准确预测熔解行为。通过将上下文序列表示与描述反应环境的显式信息相结合,该框架实现了亚度级预测精度,并将锁核酸修饰寡核苷酸的预测误差相对于最近邻热力学方法降低了高达25%。该模型还能更准确地捕捉核苷酸修饰引入的热效应,并在独立基准数据集上保持强劲性能,这些数据集的实验条件与训练时代表的条件有显著差异。我们的结果表明,学习得到的序列表示能够通过捕捉难以仅用固定参数表编码的上下文依赖效应,补充经典热力学模型。更广泛地,这项工作提供了一个可扩展的框架,用于预测跨多种化学修饰和检测条件下的寡核苷酸熔解行为,支持更可靠的分子检测设计。
英文摘要:
Oligonucleotide melting temperature is a fundamental determinant of nucleic acid hybridization and underpins the design of molecular diagnostics, polymerase chain reaction assays, and many other biotechnology applications. However, accurately predicting melting behavior remains difficult because it depends not only on sequence composition, but also on experimental conditions and chemical modifications commonly used in modern assay design. Existing thermodynamic models rely on fixed parameterizations that are often difficult to extend across diverse reaction environments and nucleotide chemistries. Here we show that a condition-aware nucleotide language model can accurately predict oligonucleotide melting behavior across diverse experimental conditions and both unmodified and chemically modified oligonucleotides. By combining contextual sequence representations with explicit information describing the reaction environment, the framework achieves sub-degree prediction accuracy and reduces prediction error for locked nucleic acid-modified oligonucleotides by up to 25% relative to nearest-neighbor thermodynamic approaches. The model also more accurately captures the thermal effects introduced by nucleotide modification and maintains strong performance on independent benchmark datasets spanning experimental conditions substantially different from those represented during training. Our results demonstrate that learned sequence representations can complement classical thermodynamic models by capturing context-dependent effects that are difficult to encode using fixed parameter tables alone. More broadly, this work provides a scalable framework for predicting oligonucleotide melting behavior across diverse chemistries and assay conditions, supporting more reliable molecular assay design.