发表机构
University of Würzburg(维尔茨堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出多尺度隐变量表示方法,为各尺度隐变量配备独立熵模型,在Kodak数据集上较VVC实现17.9%的BD-rate降低,且可与现有学习型图像压缩方法兼容。
AI 中文摘要
大多数学习型图像压缩系统依赖结合超先验的单一隐变量表示,这限制了其高效捕获不同空间尺度图像结构的能力。本研究提出一种分层隐变量表示以提升熵模型的效率,通过使用不同尺度的多个隐变量并为每个隐变量配备独立的熵模型,可更好地捕获隐变量表示的空间结构。实验表明,该方法在Kodak数据集上相比VVC实现了17.9%的BD-rate降低,验证了多尺度隐变量表示的有效性,且该方法与学习型图像压缩领域的其他进展正交,可作为现有方法的通用补充。
英文摘要
Most learned image compression systems rely on a single latent representation combined with a hyperprior, which limits their ability to efficiently capture image structure across spatial scales. In this work, we propose a hierarchical latent representation to improve the efficiency of the entropy model. By using multiple latents at different scales, each with its own entropy model, we better capture the spatial structure of the latent representation. Our experiments show that this approach achieves a 17.9% BD-rate reduction over VVC on Kodak, demonstrating the effectiveness of multi-scale latent representations. Furthermore, the approach is orthogonal to other advances in learned image compression, making it a versatile addition to existing methods.
CommentsAccepted at ICIP 2026