arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

线性时间超气泡检测的有根树框架

A rooted tree framework for linear time ultrabubble detection

Athanasios E. Zisis, Pål Sætrom

arXiv 2609.14852首次发表:更新:

发表机构

Norwegian University of Science and Technology; St. Olavs Hospital HF(挪威科技大学; 圣奥拉夫医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于有根树框架的线性时间超气泡检测方法,通过混合选择、BFS遍历及前沿节点性质,将复杂度优化至O(n+m+K),并给出六种实现与基准验证。

AI 中文摘要

泛基因组学利用图来展示物种内部或物种之间的遗传差异。在这些图中,一条路径可以代表一个基因组,而具有不同路径的区域则显示遗传变异。双色边图(biedged graph)使用黑边表示序列,灰边表示序列之间的连接。Snarl是双色边图中通过移除两条黑边而与图其余部分分离的最小次子图。超气泡(ultrabubble)是最小的无环且无尖端(tip-free)的snarl,因此它们是非常重要的变异结构,因为它们具有有限路径且没有死端。在我们之前的工作中,我们证明了在线性时间内,每个双向图(bidirected graph)都可以转换为一个有根的双色二分图(rooted biedged bipartite graph),并且在这些图中,可以使用基于最低公共祖先(LCA)的方法以$O(Kn)$时间枚举超气泡,其中$n$和$K$分别是图中的节点数和给定的snarl数。在这里,我们对我们之前基于LCA的方法进行了一系列理论和实践上的改进。首先,我们提出了一种混合方法,根据snarl的大小与图中尖端和闭环节点数量的关系,在基于LCA的方法和朴素方法之间进行选择来评估snarl。其次,通过使用我们之前论文中的理论框架,我们证明了通过遍历双色二分图的广度优先搜索(BFS)树,可以在$O(n + m + K)$时间内找到所有超气泡,其中$m$是边数。第三,我们证明了任何两个作为候选超气泡且共享一个前沿节点(frontier node)的snarl不可能是超气泡;由此产生的snarl集合是兼容的,受$n$限制,并定义了嵌套snarl的互斥族。我们将这三个结果组合成六种方法,并提供了基准测试结果,以说明上述改进如何影响识别超气泡的实际运行时间。

英文摘要

Pangenomics uses graphs to show genetic differences within or between species. In these graphs, a path can represent one genome, while regions with different paths show genetic variation. Biedged graphs use black edges for sequences and grey edges for links between them. Snarls are minimal subgraphs of a biedged graph that are separated from the rest of the graph by removing two black edges. Ultrabubbles are minimal acyclic and tip-free snarls and thus are important variant structures because they have finite paths and lack dead ends. In our previous work, we showed that in linear time every bidirected graph can be transformed to a rooted biedged bipartite one, and that in these graphs, ultrabubbles can be enumerated with a lowest common ancestor (LCA)-based method in $O(Kn)$ time, where $n$ and $K$ are the number of nodes and given snarls, respectively, of the graph. Here, we present a series of practical and theoretical improvements to our previous LCA-based approach. First, we present a hybrid method that selects between the LCA-based method and the naive approach for evaluating a snarl, depending on the size of the snarl in relation to the number of tips and cycle-closing nodes in the graph. Second, by using the theoretical framework from our previous paper, we show that all ultrabubbles can be found in $O(n + m + K)$ time, where $m$ is the number of edges, by traversing the breadth-first search (BFS) tree of the biedged bipartite graph. Third, we show that any two snarls that are candidate ultrabubbles and share a frontier node cannot be ultrabubbles; the resulting set of snarls is compatible, bound by $n$, and defines exclusive families of nested snarls. We combine these three results into six methods and present benchmarking results that illustrate how the above improvements affect practical run-times for identifying ultrabubbles.

Comments21 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑