arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeuralZip:用于快速无损压缩的可复用设置

NeuralZip: Reusable Setup for Fast Lossless Compression

Martín Bravo, Samuel Horváth, Gonzalo Navarro, Andrés Abeliuk

arXiv 2610.09916首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; University of Chile; National Center for Artificial Intelligence (CENIA)(穆罕默德·本·扎耶德人工智能大学; 智利大学; 国家人工智能中心(CENIA))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NeuralZip通过预计算并复用指数统计结构,实现快速无损压缩,设置后可转移复用,压缩速度快1.81-21.33倍,内存减少27.5%,且精确重建。

AI 中文摘要

无损压缩可以在不改变浮点数值的情况下减少模型权重的存储和移动,但重复的统计分析和代码构建增加了计算开销。我们研究了指数的统计结构是否可以一次性准备并重复使用。为此,我们引入了NeuralZip,它将具有相似指数分布的块分组,共享霍夫曼码,并选择性地使用打包指数来表示重复出现的指数元组,从而实现额外的中等压缩比。设置在后续编码之前选择这些表示,而每次编码仍然处理当前的张量值。在浮点模型检查点中,设置后的压缩比基线快1.81-21.33倍,并实现精确的逐位重建。我们表明,该设置可以预先计算并从另一个兼容架构转移,保持相似的压缩比,避免摊销设置成本的需要。因此,压缩适应是可转移和可复用的。训练检查点展示了随着权重演化的持续复用。最后,GPU实验在精确复现logits的同时,将活动内存使用量减少了高达27.5%。

英文摘要

Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios. A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values. In floating-point model checkpoints, post-setup compression is 1.81-21.33$\times$ faster than the baselines and achieves exact bit-to-bit reconstruction. We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs. Therefore, compression adaptation is transferable and reusable. Training checkpoints demonstrate continued reuse as the weights evolve. Finally, GPU experiments reduce active memory usage by up to 27.5$\%$ while reproducing the logits exactly.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑