arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机访问到LZ-End:更快且确定性的

Random Access to LZ-End: Faster and Deterministic

Itai Boneh, Paweł Gawrychowski

arXiv 2607.14923首次发表:更新:

AI 中文总结

研究针对LZ-End解析的随机访问问题,提出确定性、O(z)空间的数据结构,支持多项对数时间随机访问查询且能高效构造,查询时间优于前人,还可用于子串提取,提升了相关数据处理效率。

AI 中文摘要

长度为n的字符串的LZ-End解析是Kreft和Navarro [DCC 2010]引入的Lempel-Ziv压缩的一种变体,源于经典变体缺乏具有O(log n)访问时间的线性大小结构。虽然原始论文只能从短语边界进行有效提取,但最近Kempa和Saha [SODA 2022]确定,对于LZ-End解析由z个短语组成的字符串S,存在一种随机访问数据结构,使用O(z)空间并保证O(log⁴n·log log n)查询时间。然而,他们的证明没有产生高效的构造算法,且数据结构本质上是随机的。我们通过提供一种确定性的、O(z)空间的数据结构来解决这两个限制,该结构支持在多项对数时间内进行随机访问查询,并且可以直接从LZ-End解析在O(z log²(n/z))时间内构造。除了消除随机性并提供高效构造算法外,我们数据结构的查询时间为O(log²(n/z)),显著优于Kempa和Saha的查询时间。我们还表明我们的技术可用于支持更一般的子串提取。即,我们提出一种具有相同空间和相同构造时间的数据结构,给定两个索引i和j,在O(j - i + log²(n/z))时间内输出S[i..j]。

英文摘要

The LZ-End parsing of a length-$n$ string is a variation of Lempel-Ziv compression introduced by Kreft and Navarro [DCC 2010], motivated by the lack of a linear-size structure with $O(\log n)$ access time for the classical variant. While the original paper was only able to provide efficient extraction from the phrase boundaries, recently Kempa and Saha [SODA 2022] established that, for a string $S$ whose LZ-End parsing consists of $z$ phrases, there exists a random access data structure that uses $O(z)$ space and guarantees $O(\log^{4}n \cdot \log\log n)$ query time. However, their proof does not yield an efficient construction algorithm, and their data structure is inherently randomized. We resolve both limitations by providing a deterministic, $O(z)$-space data structure that supports random access queries in polylogarithmic time and can be constructed in $O(z\log^{2}(n/z))$ time directly from the LZ-End parsing. In addition to eliminating randomness and providing an efficient construction algorithm, the query time of our data structure is $O(\log^{2}(n/z))$, significantly improving upon the query time of Kempa and Saha. We also show that our techniques can be used to support the more general substring-extraction. Namely, we present a data structure with the same space and the same construction time that given two indices $i$ and $j$, outputs $S[i..j]$ in $O(j-i+\log^2\frac{n}{z})$ time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑