发表机构
Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对学习图像压缩中为各率失真点存独立模型的问题,提出MixCompress框架,用稀疏门控专家混合路由减轻梯度冲突,引入深度混合扩展动态扩模型容量,结合条件辅助变换,有效提升性能,建立新的帕累托前沿。
AI 中文摘要
学习图像压缩(LIC)因需要为每个率失真操作点存储独立模型而受到限制。现有可变比特率(VBR)方法试图通过密集参数调制减少开销,但强制共享主干近似不同映射会导致严重的特征纠缠。具体而言,低速率平滑梯度与高频纹理细节的保留存在固有冲突,导致性能次优。为解决此问题,我们提出MixCompress,一个基于稀疏结构专业化的统一VBR框架。稀疏门控专家混合(MoE)路由成功减轻了梯度冲突,但在固定计算预算下运行。为满足更高比特率增加的表示需求,我们引入深度混合(MoD)扩展以动态扩展模型容量。结合用于动态子带能量调制的条件辅助变换(CAT),我们的分层框架有效动态扩展容量。广泛评估表明,MixCompress不仅能匹配单独优化的单速率基线,甚至能超越它们,为高效计算图像编码建立了新的帕累托前沿。
英文摘要
Learned image compression (LIC) is bottlenecked by the need to store independent models for each rate-distortion operating point. Existing variable bit-rate (VBR) methods aim to reduce this overhead via dense parameter modulation, but forcing a shared backbone to approximate divergent mappings causes severe feature entanglement. Specifically, low-rate smoothing gradients inherently conflict with the preservation of high-frequency textural details, leading to sub-optimal performance. To resolve this, we propose MixCompress, a unified VBR framework based on sparse structural specialization. While sparsely gated Mixture-of-Experts (MoE) routing successfully mitigates gradient conflict, it operates on a fixed computational budget. To address the increased representational demands of higher bit-rates we introduce a Mixture-of-Depths (MoD) extension to dynamically scale model capacity. Combined with Conditional Auxiliary Transforms (CAT) for dynamic sub-band energy modulation, our hierarchical framework effectively dynamically scales capacity. Extensive evaluations demonstrate that MixCompress not only matches individually optimized single-rate baselines but can even surpass them, establishing a new Pareto frontier for computationally efficient image coding.
CommentsAccepted to ECCV 2026