发表机构
The Berlin Institute for Medical Systems Biology; Max Delbrück Center for Molecular Medicine; Technical University of Berlin; Humboldt University of Berlin(柏林医学系统生物学研究所; 马克斯·德尔布吕克分子医学中心; 柏林工业大学; 柏林洪堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对比三种不确定性量化方法,分析其在基因组学两类应用中的表现,明确了贝叶斯神经网络的优势,为基因组学UQ方法的应用提供了指南。
AI 中文摘要
深度学习模型已成为基因组学众多应用中的标准计算工具,但不确定性量化(UQ),尤其是该领域不同不确定性估计的可靠性,却很少受到系统关注。本研究针对基因组学应用开展深度学习模型中UQ的实证分析,在一系列实验中对比了Deep Ensembles、贝叶斯神经网络(Bayesian Neural Networks)和蒙特卡洛失活(Monte Carlo-dropout)方法,评估它们在不同场景下量化不确定性的能力,同时考虑两个基因组应用领域及模态的常见数据集特征:序列-活性模型和单细胞表达分析。我们的系统对比框架为基因组学中UQ方法的适用性与可靠性提供了指南,明确了它们在不同场景下的优势与局限性。研究表明,尽管贝叶斯神经网络存在计算劣势,但在捕捉基因组学中强类别不平衡和分布外数据导致的不确定性方面表现更优;此外,我们还展示了如何利用不确定性分数选择蛋白质-RNA相互作用中的高质量预测。
英文摘要
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications. In a series of experiments, we contrast Deep Ensembles, Bayesian Neural Networks, and Monte Carlo-dropout methods. We assess their ability to quantify uncertainty in different scenarios, accounting for common dataset characteristics in two genomic application areas and modalities: sequence-to-activity models, and single-cell expression analysis. Our systematic comparison framework provides guidelines for the applicability and reliability of UQ methods in genomics, highlighting their strengths and limitations in different scenarios. We show that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages. Moreover, we show how uncertainty scores can be used to select high-quality predictions in protein-RNA interactions.
Comments21 main pages, 42 total pages, 12 main figures, 13 supplementary figures