arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11249cs.CLcs.AIcs.ITcs.LGmath.IT

扩散到压缩:利用扩散语言模型进行无损压缩

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

Angelo Nardone, Paolo Ferragina

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对无损文本压缩的吞吐量瓶颈,首次引入扩散语言模型(DLM)替代自回归LLM,设计策略解决算法挑战,在enwik8数据集上推进了无损文本压缩的现有技术水平。

中文摘要 AI 辅助

我们研究无损文本压缩问题,其动机在于数字文本数据(包括纯文本、源代码以及XML等结构化格式)的收集与存储规模快速增长,以及基于神经语言模型的压缩技术的最新进展。特别是,近期基于大语言模型(LLM)的方法,无论是构建于符号排序流水线之上,还是与统计压缩器结合,在文本和代码上都展现出显著优于zstd、gzip或bzip等通用压缩器的压缩率。然而,这些神经方法存在严重的吞吐量限制,目前尚未具备实用价值。在无损神经文本压缩领域,我们首次引入扩散语言模型(DLMs)作为替代自回归LLM方法的推理范式。我们认为,在同一压缩框架内用DLM替代自回归LLM,可克服其每步仅处理一个符号带来的吞吐量瓶颈。但要实现这些改进,需解决将DLM应用于无损压缩时出现的算法挑战——该架构允许独立决定每次前向传播中编码符号的数量与位置。我们设计了高效且有效的策略来应对这些挑战,并在成熟的文本基准数据集enwik8上,针对基于LLM的方法和通用压缩器进行了实验评估。结果表明,新提出的基于DLM的框架推进了无损文本压缩的现有技术水平。此外,由于DLM仍是相对新兴的范式,近期在构建更强大、更高效模型方面的进展表明,该领域仍有巨大的提升空间。

英文摘要

We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression. In particular, recent LLM-based approaches, whether built on symbol-ranking pipelines or paired with a statistical compressor, have demonstrated compression ratios significantly superior to general-purpose compressors such as zstd, gzip, or bzip on text and code. However, these neural approaches suffer from severe throughput limitations, making them not yet practically usable. For the first time in the context of lossless neural text compression, we introduce Diffusion Language Models (DLMs) as an alternative inference paradigm to autoregressive LLM-based approaches. We argue that replacing autoregressive LLMs with DLMs within the same compression framework could overcome the throughput bottleneck caused by their one-symbol-per-step limitation. However, achieving these improvements requires addressing algorithmic challenges introduced by applying DLMs to lossless compression, where the architecture allows the number and positions of symbols encoded at each forward pass to be decided independently. We design efficient and effective strategies to solve these challenges and evaluate them experimentally against LLM-based and general-purpose compressors on enwik8, a well-established textual benchmark. Our results show that the newly proposed DLM-based framework advances the state of the art in lossless text compression. Moreover, as DLMs are still a relatively young paradigm, recent advances toward increasingly capable and efficient models suggest substantial room for further improvements.

补充信息

↑