FLM:用于生成式图像压缩的频率感知语言模型
FLM: Frequency-Aware Language Models for Generative Image Compression
浏览论文内容
中文总结 AI 辅助
本研究提出频率感知语言模型FLM,通过频域概率建模实现高效图像压缩,兼容有损与无损JPEG重压缩,在多数据集上较JPEG基准获显著BD-PSNR增益,性能优于传统及生成式压缩方法。
中文摘要 AI 辅助
生成模型通过利用学习到的先验,显著提升了低码率下图像有损压缩的性能上限,但生成的纹理和语义细节可能偏离源内容,从而影响图像重建的保真度。为解决这些挑战,我们提出FLM,一种频率感知语言模型,它通过频域概率建模提升压缩效率,同时保留确定性重建。在编码器端,输入图像被转换为量化的DCT系数,这些系数通过基于宏块的系数标记化被组织成离散序列;FLM随后执行下一个系数预测,以自回归方式估计用于算术编码的逐标记条件概率分布,从而生成紧凑的比特流。在解码器端,LLM与算术解码器共同恢复频域数据,之后进行逆变换以重建图像。我们还开发了特定任务的频域数据集和两阶段微调策略,使模型能够在多种码率设置下运行。FLM是一种通用压缩器,兼容有损压缩和无损JPEG重压缩框架。实验表明,FLM在率失真性能上优于传统和生成式有损压缩方法:在Kodak、Tecnick、CLIC2020数据集上,FLM较JPEG基准分别实现3.30dB、3.83dB、3.80dB的BD-PSNR增益;在语义高保真度提升和块效应抑制方面,FLM可获得更优的定性质量,且在无损重压缩任务中也被验证具有适用性与竞争力。
英文摘要
Generative models have significantly improved the performance ceiling of image lossy compression at low bitrates by exploiting learned priors. However, the generated textures and semantic details may deviate from the source content, thereby affecting the fidelity of image reconstruction. To solve these challenges, we propose FLM, a frequency-aware language model that improves compression efficiency through frequency-domain probabilistic modeling while retaining deterministic reconstruction. At the encoder, the input image is transformed into quantized DCT coefficients, which are organized into discrete sequences using macroblock-based coefficient tokenization. FLM then performs next-coefficient prediction to autoregressively estimate token-wise conditional probability distributions for arithmetic coding, thereby generating a compact bitstream. At the decoder, the LLM and arithmetic decoder jointly recover the frequency-domain data, followed by inverse transformations for image reconstruction. A task-specific frequency-domain dataset and a two-stage fine-tuning strategy are further developed to enable the model to operate across multiple bitrate settings. FLM is a versatile compressor that is compatible with both lossy compression and lossless JPEG recompression frameworks. Experiments show that FLM exceeds conventional and generative lossy compression methods in rate-distortion performance. FLM achieves BD-PSNR gains of 3.30 dB, 3.83 dB, and 3.80 dB than JPEG baseline on Kodak, Tecnick, and CLIC2020, respectively. Better qualitative quality of FLM can be achieved in improving semantically high fidelity and suppressing blocking artifacts. FLM is also validated to be applicable to the lossless recompression task with competitive performance.
发表机构
- School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)
- University of Science and Technology of China(中国科学技术大学)
- College of Intelligent Systems Science and Engineering, Harbin Engineering University(哈尔滨工程大学智能系统科学与工程学院)
- Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局信息通信研究院)
- National Tsing Hua University(国立清华大学)
机构由 AI 辅助整理,请以论文原文为准。