arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30089cs.DS

Move-rb:更快的双向 r 索引与近似模式匹配

Move-rb: Faster Bi-Directional r-indexes and Approximate Pattern Matching

Johannes Fischer, Lukas Nalbach

首次发表
浏览论文内容

中文总结 AI 辅助

Move-rb 是一种基于 Move-r 的双向 r 索引,通过优化索引操作和搜索方案,在近似模式匹配中比现有方法快 1-4 个数量级,内存更小,支持快速扩展和方向切换。

中文摘要 AI 辅助

近似模式匹配(APM)在高度重复文本上是生物信息学中的核心任务。双向 r 索引支持模式的左扩展和右扩展,从而加速基于搜索方案的 APM 算法,但现有变体——br-index 和 b-move——存在两个瓶颈:每次扩展时在游程编码 BWT 上进行 O(σ) 个字符前驱/后继查询,以及在定位时对稀疏位向量进行前驱查询以维护后缀数组区间中的值并访问 PLCP 数组。我们提出了 Move-rb,一种基于优化 r 索引 Move-r 构建的双向 r 索引。尽管 Move-rb 比 br-index 大 2 倍,但比 b-move 小最多 24%,其 APM 查询速度比 br-index 快 1 到 4 个数量级,比 b-move 快 1.9 到 10 倍,比最先进的双向(未压缩)FM 索引 columba 快最多 5.5 倍,而 columba 比 Move-rb 大 36 到 42 倍。在 APM 定位查询期间的内存使用(包括索引大小)对于汉明距离减少了 1.5 倍(最多 6.5 倍),对于编辑距离减少了 2.5 倍(最多 7.4 倍)。Move-rb 的构建速度比 br-index、b-move 和 columba 快 3 到 14 倍,同时内存使用少 19 到 141 倍。使用相对 Lempel-Ziv 编码的后缀数组的变体定位速度最多快 10 倍,同时大小增加 1.2 到 2.4 倍。我们通过优化索引操作和搜索方案 APM 算法来实现这些加速:在没有方向切换的情况下,Move-rb 以输出最优的 O(k) 时间计算所有 k 字符扩展,并以相同时间计算单字符扩展。用 O((r+\overleftarrow{r})log(n/r_min)log σ) 位增强 Move-rb 可将单字符扩展时间降至 O(log σ),其中 r 和 \overleftarrow{r} 分别是文本及其反向的 BWT 中的游程数,r_min = min(r, \overleftarrow{r})。方向切换仅增加 O(log log_ω(n/r_min)) 的额外时间。

英文摘要

Approximate pattern matching (APM) on highly repetitive texts is a central task in bioinformatics. Bi-directional r-indexes support left- and right-extension of a pattern and thereby accelerate APM algorithms based on search schemes, but existing variants -- br-index and b-move -- suffer from two bottlenecks: $O(σ)$ character-predecessor/-successor queries on the run-length-encoded BWT per extension, and predecessor queries on sparse bit vectors to maintain a value in the suffix array interval and to access the PLCP array while locating. We present Move-rb, a bi-directional r-index built on the optimized r-index Move-r. Although Move-rb is 2x larger than br-index, it is up to 24% smaller than b-move, answers APM queries 1-4 orders of magnitude faster than br-index, 1.9-10x faster than b-move and up to 5.5x faster than the state-of-the-art bi-directional (uncompressed) FM-index columba, which is 36-42x larger. Memory usage (including index size) during APM locate queries is reduced by 1.5x (up to 6.5x) for Hamming distance and 2.5x (up to 7.4x) for edit distance. Move-rb can be constructed 3-14x faster while using 19-141x less memory than br-index, b-move and columba. A variant using a relative Lempel-Ziv-encoded suffix array locates up to 10x faster while being 1.2-2.4x larger. We achieve these speedups by optimizing index operations and search scheme APM algorithms: Without a direction switch, Move-rb computes an all-$k$-character extension in output-optimal $O(k)$ time and a single-character extension in the same time. Augmenting Move-rb with $O((r+\overleftarrow{r})\log(n/r_{\min})\logσ)$ bits reduces a single-character extension to $O(\logσ)$ time, where $r$ and $\overleftarrow{r}$ are the numbers of runs in the BWT of the text and its reverse, and $r_{\min}=\min(r,\overleftarrow{r})$. A direction switch incurs only $O(\log\log_ω(n/r_{\min}))$ additional time.

发表机构

  • TU Dortmund(多特蒙德工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑