基于分层参考的生成式视频压缩
Generative Video Compression Based on Hierarchical Referencing
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于分层参考的生成式视频压缩方法GVCHR,通过分层结构优化潜在编码与生成重建,在多个基准数据集上较现有最优方法实现LPIPS和DISTS指标下50.5%、54.0%的BD-rate增益,视觉质量显著提升。
AI中文摘要:
基于扩散的生成式视频压缩已成为提升感知质量的有前景范式,其中潜在帧需被高效编码,同时作为去噪条件。然而,现有方法既未在潜在编码过程中精心设计参考结构与质量结构,也未考虑帧级质量变化对去噪过程的影响,这限制了编码效率并加剧了生成重建过程中的伪影传播。本文提出GVCHR(基于分层参考的生成式视频压缩),核心思路是将潜在帧进行分层组织,所选高质量参考帧同时有益于潜在编码与生成重建。在潜在编码阶段,GVCHR将分层参考结构与分层质量结构相结合,为作为参考被更频繁复用的低层帧分配更多比特;基于该设计,本文引入分层时间上下文挖掘技术,以利用互补的短期与长期时间上下文实现高效潜在编码。在生成重建阶段,编码侧的分层结构被整合至附加在视频扩散Transformer上的分层注意力适配器中,该适配器通过分层注意力限制每个潜在帧仅关注同层或低层参考帧,从而减少去噪过程中的伪影传播。实验在多个基准数据集上验证了GVCHR的性能,与现有最优方法相比,GVCHR在LPIPS和DISTS指标上分别实现了50.5%和54.0%的BD-rate增益,同时视觉质量也得到显著提升。
英文摘要:
Diffusion-based generative video compression has emerged as a promising paradigm to improve perceptual quality, where latent frames are required to be encoded efficiently while serving as denoising conditions. However, existing methods neither carefully design reference and quality structures during latent coding nor account for the impact of frame-level quality variation on denoising procedure, which limits coding efficiency and aggravates artifact propagation during generative reconstruction. In this paper, we propose GVCHR, Generative Video Compression based on Hierarchical Referencing. The key idea is to organize latent frames hierarchically, where the selected high-quality references benefit both latent coding and generative reconstruction. In latent coding, GVCHR couples a hierarchical reference structure with a hierarchical quality structure, assigning more bits to lower-layer frames that are reused more frequently as references. Built on this design, we introduce Hierarchical Temporal Context Mining to exploits complementary short- and long-term temporal context for effective latent coding. In generative reconstruction, the coding-side hierarchy is incorporated into a Hierarchical Attentive Adapter which is attached to a video diffusion transformer. This adapter uses hierarchical attention to restrict each latent frame to attend only to the same- or lower-layer references, thereby reducing artifact propagation during denoising. Experiments validate GVCHR on multiple benchmarks. Compared with the previous state-of-the-art method, GVCHR achieves 50.5% and 54.0% BD-rate gains in terms of LPIPS and DISTS, respectively, while also delivering clearly improved visual quality.