arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34185cs.LGcs.AI

EntroPack:任意比特率下的快速准确熵编码权重压缩

EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates

Hong Zhang, Zhongjie Duan, Yingda Chen

首次发表
浏览论文内容

中文总结 AI 辅助

EntroPack提出熵编码权重压缩方法,结合E8格量化与条件概率模型,支持任意比特率,无需校准或微调,在Z-Image-Turbo上实现比NF4更低的权重误差和更少存储。

中文摘要 AI 辅助

权重压缩有助于大型神经网络适应部署内存预算,但常见的固定宽度格式仅提供粗略的存储选择。熵编码支持更精细的速率,但实际达到的大小取决于量化权重的分布和编码开销。利用这种灵活性需要准确的速率选择和高效的权重重建以用于推理。我们提出了EntroPack,一种熵编码权重压缩器,支持任意目标比特率,无需激活校准或微调。它结合了行归一化的$E_8$格量化和格坐标的条件概率模型。采样存储估计选择量化分辨率,无需重复的全流编码。最终的坐标在可独立解码的块中进行熵编码,从而在GPU上实现快速、融合的符号解码和数值权重重建。EntroPack支持浮点和整数权重容器,如BF16、FP16、FP8和INT8,存储比特率独立于数值精度控制。在线解码增加了随权重数量增长的延迟,使该方法非常适合计算密集型工作负载,如扩散去噪和Transformer预填充。实验表明,在这些设置中,编码速度快且推理开销适中。在压缩图像生成器Z-Image-Turbo的线性层权重时,EntroPack在相当的存储速率下实现了比固定宽度格式显著更低的权重和去噪器输出误差,且去噪步骤开销适中。以每参数4比特为目标,它实现了比NF4更低的权重和去噪器输出误差,包括约24%更低的相对$L_2$权重误差,且存储更少。源代码可在该https URL获取。

英文摘要

Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility requires accurate rate selection and efficient weight reconstruction for inference. We present EntroPack, an entropy-coded weight compressor that supports arbitrary target bitrates without activation calibration or fine-tuning. It combines row-normalized $E_8$ lattice quantization with a conditional probability model of lattice coordinates. Sampled storage estimates select the quantization resolution without repeated full-stream encoding. The final coordinates are entropy-coded in independently decodable tiles, enabling fast, fused symbol decoding and numerical weight reconstruction on the GPU. EntroPack supports floating-point and integer weight containers, such as BF16, FP16, FP8, and INT8, with storage bitrate controlled independently of numerical precision. Online decoding adds latency that grows with weight count, making the method well suited to compute-intensive workloads such as diffusion denoising and Transformer prefill. Experiments demonstrate fast encoding and modest inference overhead in these settings. When compressing the linear-layer weights of the image generator Z-Image-Turbo, EntroPack achieves substantially lower weight and denoiser output errors than fixed-width formats at comparable storage rates, with modest denoising-step overhead. Targeting 4 bits per parameter, it achieves lower weight and denoiser output errors than NF4, including about 24% lower relative $L_2$ weight error, with less storage. Source code is available at https://github.com/modelscope/entropack.

↑