VANDAM:利用DNA分子先验查看核苷酸序列
VANDAM: Viewing a nucleotide sequence with DNA molecular priors
浏览论文内容
中文总结 AI 辅助
VANDAM框架通过引入DNA分子先验扩展基因组基础模型训练,提升下游任务性能,并泛化到其他分子特性。
中文摘要 AI 辅助
当代基因组基础模型(GFMs)依赖于“DNA即字符串”的范式,该范式采用掩码标记预测目标进行预训练。然而,这种抽象并未明确建模对生物功能至关重要的生化、结构和物理特性。许多分子特性可以使用已建立的生物物理模型从序列中估计,因此它们的作用不在于提供独立的模态,而在于引入训练目标可以明确利用的先验。我们提出了VANDAM,一个扩展GFMs训练以纳入DNA分子先验的框架。在自监督训练中,VANDAM从池化表示中预测区域分子特性。当功能标签可用且能够奖励保留分子先验时,局部特征还会在输入时额外注入。VANDAM通过补充基于标记的目标,一致地提升了四个架构家族和九个保留基因组任务的下游性能。探测实验进一步表明,分子先验的使用可以泛化到其他未见过的分子特性。
英文摘要
Contemporary Genomic Foundation Models (GFMs) rely on a DNA-as-a-string paradigm that employs masked token prediction objectives for pretraining. However, this abstraction does not explicitly model the biochemical, structural, and physical properties essential to biological function. Many molecular properties can be estimated from sequence using established biophysical models, so their utility lies not in providing an independent modality, but in introducing priors that training objectives can explicitly exploit. We introduce VANDAM, a framework that extends the training of GFMs with DNA molecular priors. In self-supervised training, VANDAM predicts regional molecular properties from pooled representations. When functional labels are available and can reward retaining molecular priors, local features are additionally injected at the input. VANDAM consistently improves downstream performance across four architecture families and nine held-out genomic tasks by complementing token-based objectives. Probing experiments further demonstrate that the use of molecular priors generalizes to other unseen molecular properties.
发表机构
- NVIDIA(英伟达)
- Sheba Medical Center(示巴医疗中心)
- Reichman University(莱希曼大学)
- Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院)
机构由 AI 辅助整理,请以论文原文为准。