发表机构
School of Telecommunication Engineering, Xidian University; College of Computer Science, Nankai University; School of Electronics and Information, Northwestern Polytechnical University; School of Engineering, Royal Melbourne Institute of Technology(西安电子科技大学通信工程学院; 南开大学计算机学院; 西北工业大学电子信息学院; 皇家墨尔本理工学院工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于坐标的过拟合图像编解码器,提出长程上下文外推增强熵模型(LC3EM),通过互补预测模式和自适应软模式选择,提升压缩性能,在多个基准上实现显著BD-rate增益。
AI 中文摘要
基于坐标的过拟合图像编解码器因其低解码复杂度和不依赖跨图像泛化而受到越来越多的关注。然而,代表性方法如COOL-CHIC面临固有的熵建模权衡:轻量级模型容量有限,而更具表达力的模型则因传输图像特定参数而产生额外比特率开销。受传统编解码器中预测机制的启发,我们提出了一种新的熵建模策略,引入互补预测模式并采用区域自适应软模式选择,而非依赖单一学习预测器来建模多种类型的冗余。基于此概念,我们开发了长程上下文外推增强熵模型(LC3EM),可集成到基于坐标的过拟合编解码器中。具体而言,无参数的基于邻域的线性外推模式(NLEM)补充了基于微型MLP的局部预测器,以利用长程上下文冗余和强方向性结构。设计了最小熵启发的连续模式选择策略,自适应融合这两种互补模式,仅需传输单个额外线性层的参数。此外,为缓解训练时松弛量化与实际离散量化之间的不匹配,我们引入了轻量级迭代潜在舍入细化阶段以提升压缩性能。实验表明,在多个基准上均有一致改进,尤其是在高度规则的计算机生成图像上。当与COOL-CHIC 4.0集成时,所提方法在SIQAD和API数据集上分别实现-3.43%和-7.69%的BD-rate增益。以COOL-CHIC 5.0为骨干时,相应增益分别为-2.88%和-3.15%。代码即将公开。
英文摘要
Coordinate-based overfitting image codecs have attracted increasing attention for their low decoding complexity and independence from cross-image generalization. However, representative approaches such as COOL-CHIC face an inherent entropy-modeling trade-off: lightweight models have limited capacity, while more expressive ones incur additional bitrate overhead from transmitting image-specific parameters. Inspired by the prediction mechanism in traditional codecs, we propose a new entropy-modeling strategy that introduces complementary prediction modes with region-adaptive soft mode selection, rather than relying on a single learned predictor to model diverse types of redundancy. Based on this concept, we develop a Long-Range Context Extrapolation Enhanced Entropy Model (LC3EM), which can be integrated into coordinate-based overfitting codecs. Specifically, a parameter-free Neighborhood-based Linear Extrapolation Mode (NLEM) complements the tiny MLP-based local predictor to exploit long-range contextual redundancy and strongly directional structures. A Minimum-Entropy-Inspired Continuous Mode Selection strategy is designed to adaptively fuse these two complementary modes, while requiring the transmission of only the parameters of a single additional linear layer. Moreover, to alleviate the mismatch between training-time relaxed and actual discrete quantization, we introduce a lightweight iterative latent rounding refinement stage to improve compression performance. Experiments demonstrate consistent improvements across diverse benchmarks, particularly on highly regular computer-generated images. When integrated with COOL-CHIC 4.0, the proposed method achieves BD-rate gains of -3.43\% and -7.69\% on the SIQAD and API datasets, respectively. With COOL-CHIC 5.0 as the backbone, the corresponding gains are -2.88\% and -3.15\%, respectively. The code will be made publicly available soon.