发表机构
North Carolina State University(北卡罗来纳州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出了一种新的压缩方法,通过直接从打包文本支持逆后缀数组查询,实现更快的查询时间,同时减少存储需求。
AI 中文摘要
后缀数组(SA)和逆后缀数组(ISA)是字符串算法中的基本数据结构。给定一个长度为n的整数字母表上的文本,显式存储这些结构需要Θ(n log n)位,这促使了接近信息论最小值n log σ位的压缩表示。两个标志性解决方案是FM索引,它实现了紧凑空间并支持SA和ISA查询在近对数时间内;以及压缩后缀数组(CSA),它使用O(n log σ)位并以O(log_σ^ε n)时间回答两种查询。几乎所有后续的压缩索引,包括熵压缩和重复性感知变种,都使用本质上相同的技术支持SA和ISA查询,导致相似的空间-时间折衷。这引发了一个根本问题:在相同渐近空间限制下,SA和ISA是否有相同的固有查询复杂度?我们提供证据表明这种对称性并不内在。在O(n log σ)位空间限制下,我们证明SA查询不太可能在log^{o(1)} n时间内完成,这表明改进经典的CSA界限需要重大突破。相反,我们提出了一种新的编码,支持ISA查询在O(τ + log log n / log log σ)时间内使用(1 + 1/τ + o(1))n log σ位,对于任何τ≥1。我们的主要见解是LF映射,传统上通过Burrows-Wheeler变换实现,可以适应直接从打包文本支持ISA查询,无需存储BWT,同时实现显著更快的查询时间。
英文摘要
The suffix array ($\SA$) and inverse suffix array ($\ISA$) are fundamental data structures in string algorithms. For a text $T[0 \dd n)$ over an alphabet $[0 \dd σ)$, explicitly storing either structure requires $Θ(n\log n)$ bits, whereas the text itself requires only $n\logσ$ bits. Classical compressed indexes, including the FM-index of Ferragina and Manzini [FOCS 2000] and the compressed suffix array of Grossi and Vitter [STOC 2000], reduce the space to $O(n\logσ)$ bits while supporting both $\SA$ and $\ISA$ queries efficiently. Moreover, essentially all known approaches yield nearly identical space--time trade-offs for the two operations. This raises a natural question: \emph{under the same asymptotic space bound, do $\SA$ and $\ISA$ queries have the same inherent query-time complexity?} We provide strong evidence that they do not. Achieving $\log^{o(1)} n$ query time for $\SA$ using $O(n\logσ)$ bits would require a major breakthrough in computational geometry, by a straightforward consequence of a hardness result of Chien {\it et al.} [Algorithmica 2015]. In contrast, we give a near-succinct encoding using $n\logσ+o(n\logσ)+O(n)$ bits that supports $t_{\ISA}=O!\left(τ+\frac{\log\log n}{\log\logσ}\right)$, where $τ=ω(1)$ can grow arbitrarily slowly, and $t_{\SA}=O(t_{\ISA}τ\log_σ n)$. The leading $n\logσ$ bits are used solely to store the text in packed form, enabling optimal-time substring extraction and efficient packed pattern searching. Building on Sadakane [SODA 2002], we also give a suffix-tree encoding auxiliary to the text that uses only $O(n\log\log\logσ)$ bits and supports standard operations in polylogarithmic time. For growing alphabets, this auxiliary space is asymptotically smaller than the space required to store the text itself.
CommentsTo appear in the Proceedings of the 67th IEEE Symposium on Foundations of Computer Science (FOCS 2026)