arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36357cs.AI

HyperZip:通过超网络实现个性化扩散大语言模型的高效数据压缩

HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks

  • New Jersey Institute of Technology(新泽西理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Thai Nguyen, Khang Tran, NhatHai Phan

AI总结:

针对LLM压缩中自回归解码效率低的问题,提出HyperZip框架,利用扩散LLM和多令牌预测加速压缩,并通过超网络生成数据特定更新以平衡压缩率与速度,实验证明其优于现有基线。

AI中文摘要:

大型语言模型(LLMs)在无损数据压缩方面展现出强大潜力,但现有方法受限于自回归解码的高计算成本和低吞吐量。我们提出HyperZip,一种高效且可扩展的基于LLM的压缩框架,利用基于扩散的LLM(dLLMs)与多令牌预测(MTP)来加速基于LLM的数据压缩过程。我们识别出基于扩散的压缩中存在一种权衡,即增加解码吞吐量会降低压缩率。为缓解这一权衡,HyperZip采用超网络从上下文表示生成数据特定的更新,无需昂贵的微调即可使dLLM适应目标数据,从而实现低压缩率和高吞吐量。大量实验表明,与最先进的基线相比,HyperZip在压缩率和速度之间实现了更优的权衡。

英文摘要:

Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages diffusion-based LLMs (dLLMs) with Multi-Token Prediction (MTP) to accelerate LLM-based data compression processes. We identify a trade-off in diffusion-based compression, where increasing decoding throughput degrades the compression rate. To mitigate this trade-off, HyperZip employs a hypernetwork to generate data-specific updates from a context representation, adapting the dLLM to the target data without costly fine-tuning, resulting in a low compression rate and high throughput. Extensive experiments show that HyperZip achieves a superior trade-off between compression rate and speed compared with state-of-the-art baselines.

↑