arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10188cs.DCcs.PF

GPU LZ77解码的实际序列化因素:三种解码器、三种机制及一种可消除最后一种的编码时手段

What Actually Serializes GPU LZ77 Decode: Three Decoders, Three Mechanisms, and an Encode-Time Lever That Removes the Last One

Yakiv Shavidze

中文总结 AI 辅助

该研究通过在H100上对三种解码器的测量,明确GPU LZ77解码的顺序瓶颈为解析而非复制,揭示自重叠匹配可并行,还提出编码器可消除四入口距离历史以减少顺序依赖,同时报告了格式瓶颈及反驳的10个假设。

中文摘要 AI 辅助

GPU LZ77解码的顺序部分并非该领域所认为的那样。在H100上对三种解码器架构的测量显示,解析(parse)而非复制(copy)占据了设备驻留解码时间的64%-72%;通过限制反向引用链深度(该操作可被证明,且仅产生0.006%的比率开销),延迟最多可降低2.8%,而对于文件自身的延迟峰值,由于对全部15499个块进行字节级比较显示该上限未改变涉及的181个块,因此该操作对延迟峰值无影响;自重叠匹配属于周期性填充而非依赖链,这使其可完全并行,并以比特精确的方式将匹配层提速2.75-8.42倍;最后一个真正的顺序元素——四入口距离历史(four-entry distance history)——可由编码器消除,仅产生0.540%的比率开销,使无依赖解析运行从4条命令增至706条。我们还报告了该格式面临的瓶颈:中位数匹配为7字节,对应128字节缓存行,总线效率为4.4%,而对相同数据的合并写入速度快39倍。另有一节记录了这些测量结果反驳的10个假设,包括我们自身的一个方法学错误。所有可复现的主张均带有可机器检查的记录:克隆标记版本后,无需GPU即可通过17项检查中的全部17项,无一项失败。

英文摘要

The sequential part of GPU LZ77 decode is not where the field assumes it is. Across three decoder architectures on an H100 we measure that parse, not copy, holds 64-72% of device-resident decode time; that bounding back-reference chain depth - provable, and costing 0.006% in ratio - moves latency by at most 2.8% and, for the file's own latency spike, provably by nothing at all, since a byte-level comparison of all 15,499 blocks shows the cap alters none of the 181 blocks involved; that self-overlapping matches are periodic fills rather than dependency chains, which makes them fully parallel and speeds the match layer by 2.75-8.42x bit-perfect; and that the last genuinely sequential element, a four-entry distance history, can be removed by the encoder for 0.540% of ratio, growing the dependency-free parse run from 4 commands to 706. We also report the floor the format runs into: with a median match of 7 bytes against a 128-byte cache line, bus efficiency is 4.4% and a coalesced write of the same data is 39x faster. A separate section records ten hypotheses these measurements refuted, including one methodological error of our own. Every reproducible claim carries a machine-checkable record: a fresh clone of the tagged release passes 17 of 17 checks reachable without a GPU, none failing.

补充信息

↑