AI 中文总结
该研究提出比特分配迁移框架,通过DCVC-FM训练量化步长生成模型,将神经压缩的感知重要性迁移至传统编解码器,在HEVC B~D数据集上实现显著比特率节省,且不修改核心解码语法。
AI 中文摘要
传统基于块的视频编解码器,如H.264/AVC、H.265/HEVC和H.266/VVC,依赖手工设计的率失真优化(RDO)过程,该过程主要最小化均方误差(MSE),而MSE与人类感知质量的相关性较差。虽然神经视频压缩方法可以轻松优化MS-SSIM等与感知对齐的指标,但其高计算复杂度限制了实际部署。本文提出一种新颖的比特分配迁移框架,以桥接这两种范式,增强传统视频编解码器的感知质量。具体而言,我们在神经视频压缩框架DCVC-FM内,利用感知损失训练量化步长生成模型,该模型以原始帧和运动补偿预测为输入,输出量化步长图。我们随后从该图导出块级比特率,并将其转换为传统视频编解码器的量化参数(QP)图。在HEVC B~D数据集上的实验结果表明,与标准参考软件JM-19.0、HM-16.20和VTM-23.0相比,我们的方法在MS-SSIM指标下分别实现了20.20%、8.25%和8.37%的比特率节省,利用预测帧时还能获得额外增益。我们的方法可有效将神经视频压缩模型学习到的隐式感知重要性迁移,以指导传统视频编解码器的块级比特分配,且无需修改其核心解码语法。
英文摘要
Traditional block-based video codecs, such as H.264/AVC, H.265/HEVC and H.266/VVC, rely on hand-crafted Rate-Distortion Optimization (RDO) processes that primarily minimize Mean Squared Error (MSE), which correlates poorly with human perceptual quality. While neural video compression methods can easily optimize perceptually aligned metrics like MS-SSIM, their high computational complexity limits practical deployment. This paper proposes a novel bit allocation transfer framework that bridges these two paradigms to enhance the perceptual quality of conventional video codecs. Specifically, we train a quantization step generation model using a perceptual loss within a neural video compression framework (DCVC-FM). The model takes the original frame and a motion-compensated prediction as input and outputs a quantization step map. We then derive a block-wise bit ratio from this map and convert it into a Quantization Parameter (QP) map for a traditional video codec. Experimental results on the HEVC B$\sim$D dataset demonstrate that our method achieves 20.20\%, 8.25\%, and 8.37\% bitrate savings in terms of MS-SSIM compared with the standard reference software JM-19.0, HM-16.20, and VTM-23.0, respectively, with additional gains when utilizing predicted frames. Our approach effectively transfers the implicit perceptual importance learned by neural video compression models to guide block-level bit allocation in traditional video codecs without modifying their core decoding syntax.
Comments5 pages, 4 figures