arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于字符串同步集的实用且空间高效的LZ77与LZ预压缩

Practical and Space-Efficient LZ77 and LZ Pre-Compression via String Synchronizing Sets

Jonas Ellert, Lukas Nalbach

arXiv 2609.30193首次发表:更新:

发表机构

CWI Amsterdam; TU Dortmund University(荷兰数学与计算机科学研究学会阿姆斯特丹中心; 多特蒙德工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过字符串同步集替换LZ77因子分解中的不可实现组件,首次实现实用且空间高效的LZ77及LZ预压缩算法,显著提升速度并降低内存。

AI 中文摘要

Lempel-Ziv(LZ77)因子分解将文本分解为尽可能少的 $z$ 个短语,每个短语引用一个较早的出现位置。正是这个短语数量(而非编码大小)决定了基于LZ的压缩索引的大小,而计算具有少量短语的因子分解在其构建过程中是时间和空间的瓶颈。在实践中,快速计算LZ77目前需要构建后缀数组。Ellert [SPIRE 2023] 给出了在亚线性工作空间内计算精确LZ77因子分解及其3-近似算法的算法。这些算法尚未实现,因为其中两个组件难以直接实现:一个查找表,对于实际输入会退化为长度至多二的模式;以及一个不实用的正交范围报告数据结构。我们替换了这两个组件,微调了所有剩余阶段,并获得了第一个实用实现,其运行空间接近文本大小而非后缀数组大小。在单线程上,我们的3-近似算法比经典LPF算法快12-19倍,同时内存使用减少14倍;在32线程上,即使我们的精确算法也比并行LPF快1.4--2.9倍,内存减少9倍。实际上,近似比远低于3。作为附带结果,仅将其完美短语传递给下游压缩器,得到的预压缩器在压缩比上与最先进技术[Dinklage, SEA 2026]相当,且在内存消耗和并行吞吐量方面更优。

英文摘要

The Lempel-Ziv (LZ77) factorization decomposes a text into the least possible number $z$ of phrases that each refer to an earlier occurrence. It is this phrase count, rather than the encoded size, that governs the size of LZ-based compressed indexes, and computing a factorization with few phrases is a time and space bottleneck in their construction. In practice, computing LZ77 quickly has so far required building a suffix array. Ellert [SPIRE 2023] gave algorithms that compute the exact LZ77 factorization, and a 3-approximation of it, in sublinear working space. They have remained unimplemented, because two of their components resist a direct implementation: a lookup table that degenerates to patterns of length at most two for realistic inputs, and an orthogonal range reporting data structure that is impractical. We replace both, fine-tune every remaining stage, and obtain the first practical implementation, which runs in space close to the text rather than to the suffix array. On one thread, our 3-approximation factorizes 12-19x faster than the classical LPF algorithm while using 14x less memory; on 32 threads, even our exact algorithm is 1.4--2.9x faster than parallel LPF, at 9x less memory. In practice the approximation ratio stays far below 3. As a side result, passing only its perfect phrases to a downstream compressor yields a precompressor that is on par with the state of the art [Dinklage, SEA 2026] in compression ratio, and better in memory consumption and parallel throughput.

Comments17 pages, 6 figures, 4 tables. Accepted at ALENEX 2027. Code: https://github.com/LukasNalbach/lz77-sss

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑