AI 中文总结
本研究针对系统发育网络推断的非可识别性问题,开发了计算n元组一致因子的算法及Macaulay2实现,发现五元组一致因子可识别1级网络的更多特征,为相关推断工作奠定基础。
AI 中文摘要
多种用于系统发育网络推断和非树状关系检验的统计方法,基于通过四元组一致因子(基因树中四分类单元拓扑关系的频率)评估基因组数据。这种方法虽避免了多个不良建模假设,但也导致网络根和小循环的非可识别性问题。本工作提供了一种算法及配套的Macaulay2实现,用于计算任意系统发育网络上的n元组一致因子。我们将该算法应用于五元组一致因子(总结五分类单元基因树),以探究网络多物种溯祖模型下1级网络的可识别性。结果表明,一些通过四元组无法识别的额外网络特征变得可识别。由于可识别性是任何方法进行推断的必要前提,这为未来的推断工作奠定了基础。
英文摘要
Several statistical methods of phylogenetic network inference and testing for non-tree-like relationships are based on assessing genomic data through quartet Concordance Factors, the frequencies of 4-taxon topological relationships on gene trees. While such an approach obviates making several undesirable modeling assumptions, it also results in non-identifiability issues for network roots and for small cycles. In this work, an algorithm and accompanying Macaulay2 implementation are provided for computing $n$-tet Concordance Factors on any phylogenetic network. We employ this algorithm on quintet Concordance Factors, summarizing 5-taxon gene trees, to explore identifiability of level-1 networks under the Network Multispecies Coalescent model. We show some additional network features become identifiable that are not through quartets. As identifiability is a necessary prerequisite to inference by any method, this lays a foundation for future inference work.
Comments31 pages, 11 figures, 4 tables