用于贝叶斯系统发育推断的自适应时间树转换核
An adaptive time-tree transition kernel for Bayesian phylogenetic inference
浏览论文内容
中文总结 AI 辅助
针对贝叶斯系统发育推断中树拓扑探索效率低的问题,提出自适应树转换核 subTreeLeap,其能高效探索后验树空间,且效率优于标准树转换核。
中文摘要 AI 辅助
贝叶斯系统发育与系统动力学分析非常耗时,原因在于需结合复杂模型从日益庞大的基因组数据集及关联元数据中估计关键参数。高性能计算机硬件的使用可在一定程度上缓解计算负担、显著缩短结果获取时间,但即便收敛到后验分布仍可能是漫长过程,针对大数据集的这类分析的 burn-in( burn-in 指马尔可夫链蒙特卡洛方法中前期舍弃的迭代步骤)阶段可能耗时数日甚至数周。阻碍贝叶斯系统发育推断性能的关键因素之一是树拓扑结构提议探索树空间的效率。本文提出一种名为 `subTreeLeap'(STL)的新型自适应树转换核,该方法依据可调整的半径参数,沿树中的系统发育距离路径移动以修改系统发育关系。STL 是一种通用提议方法,可用于同期或时间校准的序列数据,由于能遵守时间先后约束,尤其适用于后者。我们通过与标准树转换核下对经验数据进行长期分析得到的重复“黄金运行”结果对比,仔细评估了 STL 对后验树空间探索的收敛性与统计混合度的影响。研究发现,STL 可成功探索与标准核相同的后验树空间,但通常效率更高。我们还讨论了 STL 的局限性及未来潜在改进方向,这些改进或可大幅提升贝叶斯系统发育推断的获取速度。
英文摘要
Bayesian phylogenetic and phylodynamic analyses can be very time-consuming, owing to the combination of complex models that are used to estimate key parameters from increasingly large genomic data sets and their associated metadata. The use of high-performance computer hardware can -- to a certain extent -- alleviate the computational burden and markedly decrease the time to results. Still, even converging to the posterior can be a lengthy endeavour, with the burn-in aspect of such analyses potentially taking days or even weeks for large data sets. One of the key aspects that hampers performance in Bayesian phylogenetic inference is the efficiency with which tree topology proposals explore tree space. We here propose a novel adaptive tree transition kernel, which we call `subTreeLeap' (STL), which involves modifying the phylogeny by walking along patristic distance paths in the tree according to an adaptable radius parameter. STL is a general proposal, which can be used with contemporaneous or time-calibrated sequence data, being particularly suited to the latter due to respecting temporal precedence constraints. We carefully assess its impact on convergence and statistical mixing of the exploration of posterior tree space, by comparison to replicate ``golden runs'' obtained from lengthy analyses of empirical data under standard tree transition kernels. We find that STL successfully explores the same posterior tree space as standard kernels, but often does so in a more efficient manner. We discuss limitations as well as future potential improvements to STL that could substantially increase the speed at which Bayesian phylogenetic inferences are obtained.
发表机构
- KU Leuven(荷语鲁汶大学)
- Fred Hutchinson Cancer Center(弗雷德·哈钦森癌症中心)
- University of California, Los Angeles(加州大学洛杉矶分校)
- David Geffen School of Medicine at UCLA(加州大学洛杉矶分校大卫格芬医学院)
- University of Edinburgh(爱丁堡大学)
- Getulio Vargas Foundation(热图利奥·瓦加斯基金会)
机构由 AI 辅助整理,请以论文原文为准。