发表机构
School of Engineering Science, Simon Fraser University(工程科学学院,西蒙·弗雷泽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对H.264视频编码中优化量化参数难的问题,提出可微代理学习方法,基于可变速率学习压缩模型,通过软索引机制使代理可微,经训练近似H.264率失真行为,所构建框架在相关任务中显著改善率任务权衡。
AI 中文摘要
在过去二十年中,H.264因其相对简单、高效以及软硬件实现广泛而成为最常用的视频编码格式。然而,由于标准视频编解码器的不可微性质,针对特定目标(如感知质量或机器视觉任务)优化编解码器参数(如量化参数(QP))具有挑战性。虽然最近可微代理已被用于在标准编解码器周围进行基于梯度的优化,但其对目标编解码器的保真度很少被明确表征。本文提出一种用于H.264帧内编解码器的可微代理学习方法,以实现自适应量化控制。该代理基于变速率学习压缩模型构建,通过软索引机制使其相对于编解码器QP可微。然后在两种量化设置下进行训练,以近似H.264的率失真行为。使用训练好的代理,开发了一个基于代理的自适应量化(AQ)框架用于感知优化和机器视觉任务。实验结果表明,所提出的代理紧密近似H.264帧内编解码器的率失真行为。基于代理的AQ框架在固定QP的H.264基线之上持续改善率任务权衡,在语义分割中实现高达17.12%的BD率降低,在MS-SSIM中实现15.30%的BD率降低。
英文摘要
H.264 has been the most widely used video coding format for the past two decades due to its relative simplicity, efficiency, and wide availability of software and hardware implementations. However, optimizing codec parameters such as the quantization parameter (QP) for specific objectives (e.g., perceptual quality or machine vision tasks) is challenging due to the non-differentiable nature of standard video codecs. While differentiable proxies have recently been used to enable gradient-based optimization around standard codecs, their fidelity to the target codec is rarely explicitly characterized. In this paper, we propose a differentiable proxy learning method for H.264 intra codec to enable adaptive quantization control. Built upon a variable-rate learned compression model, the proposed proxy is made differentiable with respect to codec QP through a soft-indexing mechanism. It is then trained to approximate the rate-distortion behavior of H.264 under two quantization settings: global-QP, which uses one QP per image, and spatial-QP, which assigns QPs at the macroblock level. Using the frozen trained proxy, we develop a proxy-based adaptive quantization (AQ) framework for both perceptual optimization and machine vision tasks. Experimental results demonstrate that the proposed proxies closely approximate the rate-distortion behavior of H.264 intra codec. The resulting proxy-based AQ framework consistently improves rate-task trade-offs over fixed-QP H.264 baselines, achieving BD-rate reduction of up to 17.12% for semantic segmentation and 15.30% for MS-SSIM.
CommentsAccepted by IEEE MIPR 2026