arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

紧凑空间下的常数时间逆后缀数组查询与后缀数组索引的亚线性时间构造

Constant-Time Inverse Suffix Array Queries in Compact Space and Sublinear-Time Construction of Suffix Array Indexes

Dominik Kempa, Tomasz Kociumaka

arXiv 2608.19123首次发表:更新:

AI 中文总结

该研究提出最优空间的常数时间逆后缀数组结构,在字RAM中确定性构造该结构及两类后缀数组,匹配二元文本的时间-空间下界,可提升字典匹配算法效率。

AI 中文摘要

对于文本 $T\in[0..\sigma)^n$(其中 $2\leq\sigma\leq n$),其后缀数组按字典序排列后缀起始位置,而逆后缀数组将每个位置映射到对应后缀的秩。自2000年压缩后缀数组与FM-index出现以来,核心目标是在 $O(n\log\sigma)$ 比特空间内支持这两种查询。Thankachan近期将逆后缀数组查询时间降至 $O(\log\log n/\log\log\sigma)$,但常数时间的实现仍未解决。我们提出首个兼具最优空间与查询时间的逆后缀数组结构:空间为 $O(n\log\sigma)$ 比特,查询时间为 $O(1)$。对于二元文本,该结构无条件分离了确定性结构下的两种查询,因为在单元探测模型中,每个 $O(n)$ 比特的后缀数组结构(单元大小为 $\Theta(\log n)$ 比特)的最坏情况查询时间为 $\Omega(\log\log n/\log\log\log n)$。构造是另一项挑战:线性时间构造可能需要 $\Theta(\log_{\sigma} n)$ 倍于读取输入或写入紧凑索引的时间。此前仅有一种支持两种查询的索引已知存在亚线性构造。在字长为 $\Theta(\log n)$ 比特的字RAM中,我们从打包文本确定性构造了该新结构及两类后缀数组,时间为 $O(n\min(1,\log\sigma/\sqrt{\log n}))$。对于 $B\geq2$,第一类使用 $O(n\log\sigma(1+\log_B\log_{\sigma} n))$ 比特空间,查询时间为 $O(B(1+\log_B\log_{\sigma} n))$;第二类使用 $O(Bn\log\sigma(1+\log_B\log_{\sigma} n))$ 比特空间,查询时间为 $O(1+\log_B\log_{\sigma} n)$。每类的峰值预处理空间均受限于其索引大小。对于二元文本,当 $B\geq(\log\log n)^{\Omega(1)}$ 时,第二类匹配确定性单元探测时间-空间下界;在最慢查询场景之外,若将确定性构造时间改进为 $o(n/\sqrt{\log n})$,将得到同样快速的字典匹配算法。

英文摘要

For a text $T\in[0..σ)^n$ with $2\leqσ\leq n$, its suffix array orders the suffix starting positions lexicographically, while its inverse suffix array maps each position to its suffix's rank. Since compressed suffix arrays and FM-indexes appeared in 2000, a central goal has been to support both queries in $O(n\logσ)$ bits. Thankachan recently reduced inverse suffix array query time to $O(\log\log n/\log\logσ)$, but constant time remained open. We give the first inverse suffix array structure with optimal space and query time: $O(n\logσ)$ bits and $O(1)$ time. For binary texts, this unconditionally separates the two queries for deterministic structures, since every $O(n)$-bit suffix array structure in the cell-probe model with $Θ(\log n)$-bit cells has worst-case query time $Ω(\log\log n/\log\log\log n)$. Construction is a second challenge: linear time can take $Θ(\log_σ n)$ times as long as reading the input or writing a compact index. Previously, sublinear construction was known for only one such index supporting both queries. In the word RAM with $Θ(\log n)$-bit words, we deterministically construct the new structure and two suffix array families from the packed text in $O(n\min(1,\logσ/\sqrt{\log n}))$ time. For $B\geq2$, the first family uses $O(n\logσ(1+\log_B\log_σn))$ bits and has query time $O(B(1+\log_B\log_σn))$, whereas the second uses $O(Bn\logσ(1+\log_B\log_σn))$ bits and has query time $O(1+\log_B\log_σn)$. Each has peak preprocessing space bounded by its index size. For binary texts, the second family matches the deterministic cell-probe time-space lower bound whenever $B\geq(\log\log n)^{Ω(1)}$, and, outside the slowest-query regimes, improving the deterministic construction time to $o(n/\sqrt{\log n})$ would yield an equally fast Dictionary Matching algorithm.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑