后缀数组查询的单元探测下界与复杂度保持归约
Cell-Probe Lower Bounds and Complexity-Preserving Reductions for Suffix Array Queries
浏览论文内容
中文总结 AI 辅助
本文证明了后缀数组查询的首个单元探测时空下界,解答了Grossi与Vitter的25年问题,还给出后缀数组查询与短串前缀查询间无附加项的复杂度保持等价归约。
中文摘要 AI 辅助
对于大小为σ的字母表上长度为n的文本T,其后缀数组按字典序列出T所有后缀的起始位置,逆后缀数组给出每个位置起始后缀的字典序排名。自2000年FM-index和压缩后缀数组问世以来,两类查询均可在任意常数ε>0下,以O((log_σ n)^ε)的查询时间、O(n log σ)的空间实现,但此前尚无关于后缀数组查询的非平凡时空下界。本文给出首个此类下界:在字长为Θ(log n)的单元探测模型中,所有用于处理长度不超过n的二进制串后缀数组查询的S位数据结构,其查询时间为Ω(log log n / log((S/n) log log n))。由此可得:所有使用O(n (log log n)^O(1))位空间的结构,查询时间为Ω(log log n / log log log n);而常数查询时间则需要Ω(n log^ε n)位空间(对某常数ε>0),特别地,不存在O(n)位的二进制文本后缀数组表示能支持常数时间查询,解答了Grossi和Vitter提出的25年历史问题。此外,本文还给出后缀数组访问与短串上更简单前缀查询之间的精确复杂度保持等价关系:对任意2≤σ≤n,后缀数组查询等价于前缀选择查询,逆后缀数组查询等价于前缀特殊排名查询;双向归约在四种标准度量(空间、查询时间、预处理时间、预处理空间)上均保持常数因子,且不同于此前归约,无附加的O(log log n)查询时间项,因此对应的前缀查询问题可在四种度量上无渐近损失地刻画后缀数组与逆后缀数组访问。
英文摘要
For a text $T$ of length $n$ over an alphabet of size $σ$, its suffix array lists the starting positions of the suffixes of $T$ in lexicographic order, and its inverse suffix array gives the lexicographic rank of the suffix starting at each position. Since the introduction of the FM-index and the compressed suffix array in 2000, both queries have been supported in $O((\log_σn)^ε)$ time using $O(n\logσ)$ bits, for any constant $ε>0$. Yet no nontrivial time-space lower bound for suffix-array queries was known. We give the first such lower bound. Specifically, we show that, in the cell-probe model with $Θ(\log n)$-bit words, every $S$-bit data structure answering suffix-array queries on binary strings of length at most $n$ has query time $Ω(\log\log n/\log((S/n)\log\log n))$. Consequently, every structure using $O(n(\log\log n)^{O(1)})$ bits requires $Ω(\log\log n/\log\log\log n)$ query time, while constant query time requires $Ω(n\log^εn)$ bits for some constant $ε>0$. In particular, no $O(n)$-bit suffix-array representation for binary texts supports constant-time queries, answering the 25-year-old question of Grossi and Vitter. We also give exact complexity-preserving equivalences between suffix-array access and simpler prefix queries on short strings. For every $2\leqσ\leq n$, suffix-array queries are equivalent to prefix-select queries, and inverse-suffix-array queries are equivalent to prefix-special-rank queries. The reductions in both directions preserve all four standard measures up to constant factors: space, query time, preprocessing time, and preprocessing space. Unlike previous reductions, they incur no additive $O(\log\log n)$ query-time term. Thus, the corresponding prefix-query problems capture suffix-array and inverse-suffix-array access without asymptotic loss in any of the four measures.