arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37328eess.IV

CASR:面向多功能视频编码的内容自适应神经超分辨率后置滤波器,基于低秩过拟合

CASR: Content-Adaptive Neural Super-Resolution Post-Filter for Versatile Video Coding via Low-Rank Overfitting

  • Nokia(诺基亚)
  • Aalto University(阿尔托大学)
  • Tampere University(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

Khoa Pham-Dinh, Francesco Cricri, Maria Santamaria, Honglei Zhang, Hamed R. Tavakoli, Moncef Gabbouj, Juho Kannala, Miska M. Hannuksela

中文总结 AI 辅助

本文提出CASR,一种基于编码端低秩过拟合的内容自适应超分辨率后置滤波器,用于VVC,通过冻结预训练网络并微调LoRA矩阵,在较小信令开销下实现显著BD-rate节省。

中文摘要 AI 辅助

将超分辨率作为视频编解码器之后的后处理步骤,允许在编码端以降低的空间分辨率对视频进行编码,并在解码端重建并上采样至原始分辨率。这种方式在不修改核心编码架构的情况下,降低了所需比特率并提高了重建帧的质量。然而,通用超分辨率模型通常离线训练,缺乏对视频编解码器产生的多样内容特征和压缩伪影的充分适应性,这限制了其在测试时的有效性。为解决这一局限,本文提出CASR(内容自适应超分辨率),一种面向多功能视频编码(VVC)的内容自适应超分辨率后置滤波器框架,基于编码端对每个输入序列的过拟合。为限制传输内容自适应信号(即权重更新)所需的比特率开销,采用了低秩自适应(LoRA)技术。该方法冻结预训练超分辨率网络的卷积核,仅微调附加在选定卷积层上的轻量级秩r矩阵,使用VVC解码帧和测试序列的量化参数(QP)图作为监督。所得低秩更新采用MPEG神经网络压缩与表示(NNR)标准进行压缩。在JVET通用测试条件(CTC)A1和A2类序列上的实验表明,基于LoRA的内容自适应在较小信令开销下,相较于非自适应超分辨率后置滤波器,提供了比特率节省。与VVC测试模型(VTM21)相比,所提方法在随机接入下实现了BD-rate节省-10.93%(Y)、-15.39%(U)、-24.41%(V),在全帧内下实现了-13.43%(Y)、-5.75%(U)、-22.94%(V)。对LoRA秩r的消融进一步表明,r=4在编码增益和信令开销之间提供了最佳权衡。

英文摘要

The use of super-resolution as a post-processing step following a video codec allows videos to be encoded at reduced spatial resolution at the encoder side and to be reconstructed and upsampled to the original resolution at the decoder side. In this way, the required bitrate is reduced and the quality of the reconstructed frames is improved without modifying the core coding architecture. However, generic SR models are typically trained offline and lack sufficient adaptability to the diverse content characteristics and compression artifacts produced by video codecs, which limits their effectiveness during test time. To address the limitation, this paper proposes CASR (Content-Adaptive Super-Resolution), a content-adaptive SR post-filter framework for Versatile Video Coding (VVC), based on encoder-side overfitting on each input sequence. In order to limit the bitrate overhead required for signalling the content adaptation signal, i.e. the weight-update, Low-Rank Adaptation (LoRA) is leveraged. The method freezes the convolution kernels of a pretrained SR network and fine-tunes only lightweight rank-r matrices attached to selected convolution layers, using VVC decoded frames and quantization-parameter (QP) maps of test sequences as supervision. The resulting low-rank update is compressed with the MPEG Neural Network Compression and Representation (NNR) standard. Experiments on the JVET common test conditions (CTC) class A1 and A2 sequences indicate that LoRA-based content adaptation provides bitrate savings over a non-adapted SR post-filter at a small signalling cost. Compared with the VVC Test Model (VTM21), the proposed method achieves BD-rate savings of -10.93% (Y), -15.39% (U), -24.41% (V) under random access and -13.43% (Y), -5.75% (U), -22.94% (V) under all-intra. An ablation of the LoRA rank r further shows that r=4 provides the best trade-off between coding gain and signalling cost.

补充信息

↑