arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合FM索引与压缩后缀数组:在1阶秩编码上进行反向搜索

Mixing FM-indexes and CSAs: backward search over an order-1 rank encoding

Travis Gagie

arXiv 2609.40299首次发表:更新:

发表机构

Center for Biotechnology and Bioengineering (CeBiB)(生物技术生物工程中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对FM索引和CSA在字母表增大时的性能差异,提出一种基于1阶秩编码的混合索引,通过反向搜索结合单步CSA恢复信息,在合成数据上速度最快,但空间开销较大。

AI 中文摘要

FM索引和压缩后缀数组(CSAs)通常被视为可互换的,但随着字母表的增大,它们的行为会有所不同。FM索引的每一步大约在小波树的每一层产生一次缓存未命中,因此随着字母表大小的增加,它会变得更慢。CSA的每一步是一次二分搜索,其搜索范围随着字符变得越来越稀有而缩小。我们描述了一种简单的混合方法。文本中的每个字符都被替换为其在紧随前一个字符之后的字符中出现的频率的秩。我们在这个编码上进行反向搜索,该编码基于一个小型且偏斜的字母表,并通过在称为ΨE的数组上进行一次类似CSA的步骤,恢复编码丢失的一条信息(模式的首个字符)。计数是精确的,定位通过标准后缀数组采样实现。在合成重复数据上的原型表明,在中间字母表大小且含1%噪声的情况下,该混合方法是所尝试索引中最快的,但即使其紧凑版本也比原始文本的压缩游程CSA或FM索引大1.7到3.9倍,因为编码和ΨE一起比原始Burrows-Wheeler变换具有更多的游程。在真实数据(如解析树和最小化器摘要)上情况是否会改变,是主要待解决的问题。

英文摘要

FM-indexes and compressed suffix arrays (CSAs) are often treated as interchangeable, but they behave differently as the alphabet grows. An FM-index step costs about one cache miss per level of a wavelet tree, so it gets slower with the alphabet size. A CSA step is a binary search whose range shrinks as characters get rarer. We describe a simple hybrid. Each character of the text is replaced by the rank of its frequency among the characters that follow the previous character. We backward-search on this encoding, which is over a small, skewed alphabet, and recover the one piece of information the encoding loses (the first character of the pattern) with a single CSA-like step on an array we call $\PsiE$. Counting is exact, and locating works with standard suffix-array sampling. A prototype on synthetic repetitive data shows that the hybrid is the fastest of the indexes we tried at intermediate alphabet sizes with 1\% noise, but even its compact version is 1.7 to 3.9 times larger than a compressed run-length CSA or FM-index of the original text, because the encoding and $\PsiE$ together have more runs than the original Burrows--Wheeler transform. Whether that changes on real data, such as parses and minimizer digests, is the main open question.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑