arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进的LZ77压缩:匹配长度相关的滑动窗口

Improved LZ77 Compression with Match-Length-Dependent Sliding Windows

Yingquan, Wu

arXiv 2610.06530首次发表:更新:

发表机构

Tenafe Inc.(Tenafe公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出WLZ编码器族,滑动窗口大小随匹配长度变化,通过解析转移界和校准窗口调度,在保持普适性的同时降低冗余度,并改进熵率估计。

AI 中文摘要

我们设计并分析了WLZ,一族LZ77编码器,其滑动窗口大小取决于匹配长度:短匹配使用较小的窗口和较短的距离字段,而长匹配保留对远处重复的访问。设$B:=\log_2 W$为最大窗口$W$。一个解析转移界同时考虑了窗口限制和错过的匹配,保持了已有的收敛性和有限输入极小极大阶,而一个更精细的块界证明了当全窗口匹配从长度$o(B)$开始时,精确贪心WLZ的平稳遍历普适性。一个校准的窗口调度相对于单窗口基线从不增加最小令牌成本,并在指定的iid源上节省$\Omega(BW^{-\beta})$的速率,其中$\beta>2$。剩余的收益来自令牌编码而非窗口。在均匀iid源上,一个增长的短语上限与一个廉价长度符号将冗余度从原始1977年固定字段格式(无论其使用何种短语上限)的$\Omega(\log B/B)$改进到$O(\log\log B/B)$。将单符号游程$(r,1)$重新编码为$(1,r)$,其中对数成本计数$r$可能超过匹配上限,将具有$O(n^\alpha)$个游程($0\le\alpha<1$)的输入编码为$O(n^\alpha\log n)$比特,而该格式在短语上限$\Theta(\log W)$且$W=o(n)$时则需要$\Omega(n)$比特。在全局$p$-周期输入上,对具有最近距离平局的受限解析的字段进行霍夫曼编码,将大文件速率从固定宽度编码下的$\Theta(B/W)$降低到至多$3/W$,包括表和帧。最后,在完整的WLZ码中选择产生一个熵率估计器,该估计器几乎必然且在均值意义下一致,并对固定的非均匀iid源具有有限数据误差界。

英文摘要

We devise and analyze WLZ, a family of LZ77 encoders whose sliding-window sizes depend on match length: short matches use smaller windows and shorter distance fields, while long matches retain access to distant repetitions. Write $B:=\log_2 W$ for the maximum window $W$. A parsing-transfer bound charges both window restrictions and missed matches, preserving established convergence and finite-input minimax orders, and a sharper block charge proves stationary-ergodic universality of exact greedy WLZ when full-window matching begins at length $o(B)$. A calibrated window schedule never increases minimum token cost relative to a single-window baseline and saves $Ω(BW^{-β})$ in rate on specified iid sources, with $β>2$. The remaining gains come from the token code rather than the windows. On a uniform iid source, a growing phrase cap with a cheap length symbol improves redundancy from the $Ω(\log B/B)$ of the original 1977 fixed-field format, whatever phrase cap it uses, to $O(\log\log B/B)$. Recoding a single-symbol run $(r,1)$ as $(1,r)$, with a logarithmic-cost count $r$ that may exceed the match cap, codes inputs with $O(n^α)$ runs, $0\leα<1$, in $O(n^α\log n)$ bits, versus $Ω(n)$ for that format with phrase cap $Θ(\log W)$ and $W=o(n)$. On globally $p$-periodic inputs, Huffman coding the fields of a capped parse with nearest-distance ties reduces the large-file rate from $Θ(B/W)$ under fixed-width coding to at most $3/W$, including tables and framing. Finally, selection among complete WLZ codes yields an entropy-rate estimator consistent almost surely and in mean, with finite-data error bounds for fixed nonuniform iid sources.

Comments53 pages. Submitted to IEEE Transactions on Information Theory

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑