arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39315cs.CV

重新思考极低比特率下的生成式图像压缩

Rethinking Generative Image Compression at Extremely Low Bitrates

发表机构中国科学技术大学
查看机构详情
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出RAE-CoD,一种在表示自编码器空间中的压缩导向扩散模型,通过压缩与源表示直接对齐,在极低比特率(如256×256图像仅16比特)下避免语义崩溃,实现优雅过渡,并在MSCOCO-30K上显著优于现有方法。

中文摘要 AI 辅助

生成式图像压缩在低比特率下能产生视觉上合理的重建结果,然而当比特率趋近于零时,其行为在很大程度上仍未被探索。当被压缩到低于正常操作比特率时,代表性编解码器会发生语义崩溃:它们并非优雅地丢失源特定细节,而是产生畸形或无法识别的内容。我们的分析识别出两个因素。随着比特率降低,重建损失在梯度和视觉结果上与语义目标越来越冲突,而基于像素空间和重建导向的VAE扩散模型在语义保持方面效率降低。在这些发现的指导下,我们引入了RAE-CoD,一种在表示自编码器(RAE)空间中构建的压缩导向扩散(CoD)模型,通过压缩表示与源表示的直接对齐,在仅用16比特的情况下,为256×256图像保留可识别、自然结构化的内容。我们使用五个视觉基础模型(VFM)和一个盲视觉语言模型协议评估该框架。在MSCOCO-30K上,RAE-CoD在所有评估中脱颖而出。在0.001-0.008 bpp下,相对于最佳竞争者,它降低了至少25.7%的相对VFM特征MSE和69.1%的Fréchet距离比率。同时,重建的语义可识别性和质量几乎保持不变,而源一致性平滑下降,将突然的语义崩溃替换为向无条件生成的优雅过渡。代码将在https://this URL发布。

英文摘要

Generative image compression produces visually plausible reconstructions at low bitrates, yet their behavior as the rate approaches zero remains largely unexplored. When pushed below normal operating rates, representative codecs undergo semantic collapse: rather than gracefully losing source-specific detail, they produce malformed or unrecognizable content. Our analysis identifies two factors. As the bitrate decreases, reconstruction losses increasingly conflict with semantic objectives on gradients and visual results, while pixel-space and reconstruction-oriented VAE diffusion models become less efficient on semantic preservation. Guided by these findings, we introduce RAE-CoD, a compression-oriented diffusion (CoD) built in a representation autoencoder (RAE) space with direct alignment between compressed and source representations, preserving recognizable, naturally structured content for a $256\times256$ image with as few as 16 bits. We evaluate this framework using five vision foundation models (VFM) and a blinded vision-language model protocol. On MSCOCO-30K, RAE-CoD stands out from all evaluation. At 0.001-0.008 bpp, it reduces relative VFM feature MSE and Fréchet Distance ratio by at least 25.7% and 69.1% over the best competitors. Meanwhile, semantic recognizability and quality of the reconstructions remain nearly constant while source consistency falls smoothly, replacing abrupt semantic collapse with a graceful transition toward unconditional generation. Code will be released at https://github.com/LuizScarlet/RAE-CoD.

↑