发表机构
University of Bristol; Bytedance Inc., San Diego, USA(布里斯托大学; 字节跳动公司(圣迭戈))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出图启发的隐式神经表示G-NeRV,通过时间邻域消息传递和记忆库机制显式利用时间冗余,在UVG数据集上BD-rate比NVRC和VVC VTM分别降低8.86%和14.68%。
AI 中文摘要
隐式神经表示(INR)为视频压缩提供了一种紧凑且内容自适应的范式,通常通过共享网络参数和帧索引嵌入来表示视频。与传统或基于自编码器的编解码器相比,这些方法以隐式方式利用视频中的时间冗余,这可能导致次优的压缩性能。在本文中,我们提出了G-NeRV,一种图启发的INR,它在隐式潜在空间中显式地改善时间冗余的利用。受信息论中总相关原理的启发,我们在帧嵌入上构建时间邻域,并通过自适应门控执行消息传递以聚合来自相邻帧的可重用信息,该门控控制相邻信息的注入。受传统视频编码中参考帧缓冲区的启发,我们进一步设计了记忆库机制,以便在INR训练中随机帧索引采样下实现高效的时间邻域检索。这种新的表示模型已集成到先进的表示压缩框架中,并与现有的传统和神经视频编解码器进行了比较。结果表明,G-NeRV编解码器在UVG数据集上以PSNR衡量,分别比最先进的基于INR的编解码器NVRC和最新的标准视频编解码器VVC VTM高出8.86%和14.68%(以BD-rate计)。
英文摘要
Implicit Neural Representations (INR) provide a compact and content-adaptive paradigm for video compression, typically representing a video through shared network parameters and frame-indexed embeddings. Compared to conventional or autoencoder-based codecs, these approaches exploit temporal redundancy within videos in an implicit manner, which potentially results in sub-optimal compression performance. In this paper, we propose G-NeRV, a graph-inspired INR that explicitly improves temporal redundancy exploitation in the implicit latent space. Motivated by the total correlation principles in information theory, we construct a temporal neighborhood over frame embeddings and perform message passing to aggregate reusable information from neighboring frames through an adaptive gate controlling the injection of neighboring information. Inspired by the reference frame buffer in conventional video coding, a memory bank mechanism has been further designed to enable efficient temporal-neighbor retrieval under random frame-index sampling in INR training. This new representation model has been integrated into an advanced representation compression framework and compared with existing conventional and neural video codecs. The results show that the G-NeRV codec outperforms the state-of-the-art INR-based codec, NVRC, and the latest standard video codec, VVC VTM, by 8.86\% and 14.68\% (in BD-rate), respectively, measured by PSNR on the UVG dataset.
Comments12 pages, 6 Figures