发表机构
University of Pisa; Scuola Superiore Sant’Anna(比萨大学; 圣安娜高等学校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于大语言模型的源代码无损压缩问题,提出两种阈值符号排序变体,将预测限制在前\(T\)个排名,通过通用压缩器联合压缩阈值外符号。实验表明该方法在压缩率和吞吐量上优于此前方法,在压缩速度频谱中提供新权衡点。
AI 中文摘要
我们研究了源代码无损压缩问题,受大规模软件存档(如软件遗产)的存储需求驱动。通用压缩器在压缩率和速度间有良好权衡,但未充分利用源代码固有规律。近期方法在香农符号排序框架内利用大语言模型,虽能有效减少空间,但吞吐量显著下降。本文引入基于大语言模型的压缩器,部署两种新颖符号排序变体,将预测限制在前\(T\)个排名(\(T = 1\)或\(63\)),阈值外符号作为例外存储并与排名流通过通用压缩器联合压缩。我们对30个大语言模型进行了首次大规模基于大语言模型的源代码压缩评估。我们的\(T\)限制方法在压缩率(相对提高达37%)和压缩吞吐量(快40%)上均优于先前基于大语言模型的压缩器。与通用压缩器相比,我们获得高达82%的相对压缩增益,但速度较低,在压缩速度频谱中提供了新的权衡点。我们还表明,这些增益在源代码上比在自然语言上更强,这表明源代码揭示了大语言模型捕获但通用基于精确匹配的压缩器错过的规律。最后我们对提供理论和实践研究途径的开放问题进行了评论。
英文摘要
We study the problem of lossless compression of source code, motivated by the storage demands of large-scale software archives, such as Software Heritage (https://www.softwareheritage.org/). General-purpose compressors (e.g., zstd, bzip2) offer a good trade-off between compression ratio and speed, but fail to exploit all special regularities inherent in source code. Recent approaches leverage Large Language Models (LLMs) within Shannon's symbol-ranking framework, relying on a scheme in which the predicted rank can grow arbitrarily. While effective at reducing space, this setting incurs significant throughput degradation, and leaves open the question whether it is necessary to explicitly encode all ranks. In this work, we introduce LLM-based compressors deploying two novel symbol-ranking variants that bound predictions to the top-$T$ ranks ($T=1$ or $63$), with out-of-threshold symbols stored as exceptions and compressed jointly with the rank stream via general-purpose compressors. We conduct the first large-scale evaluation of LLM-based source code compression across 30 LLMs, including general-domain, code-specialized, and quantized models. Our $T$-bounded approach outperforms prior LLM-based compressors both in compression ratio (up to 37% relative improvement) and compression throughput (40% faster). Compared to general-purpose compressors (e.g., zstd, bzip2), we obtain up to 82% relative compression gain but at a lower speed, thus offering a new trade-off point in the compression-speed spectrum. We also show that these gains are stronger on source code than on natural language, suggesting an interesting indication, namely that source code exposes regularities captured by LLMs but missed by general-purpose exact-match-based compressors. We conclude by commenting on open problems that offer theoretical and practical avenues of research.
CommentsThis is the version of the paper submitted and then accepted for presentation at the 24th International Conference of the Italian Association for Artificial Intelligence (AIxIA 2026), Perugia, Italy, 6--9 October 2026. The paper will appear in the proceedings, published by Springer Verlag in the Lecture Notes in Artificial Intelligence (LNAI) series