arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34725eess.IV

Gen2-VC:解锁视频压缩的生成先验

Gen2-VC: Unlocking Generative Priors for Video Compression

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yinhuan Huang, Jingkai Ying, Pu Chen, Zhijin Qin

AI总结:

Gen2-VC通过冻结编解码器和生成骨干,仅用轻量LoRA适配器精化重建,实现低比特率下优于VTM-23.0的失真-感知权衡,且适配器可跨编解码器泛化。

AI中文摘要:

在严格的比特率约束下,现有视频编解码器难以在源保真度和感知真实感之间取得平衡。面向失真的编解码器常常过度平滑细节,而生成式编解码器则可能引入内容和结构偏差,并依赖特定于编解码器的设计。这引发了一个问题:现有的编解码器能否通过简单、可复用的适配实现更好的失真-感知权衡?我们的洞察是,原生编解码器的重建提供了一个共享接口,通过该接口,生成式精化被锚定到源内容,同时与编解码器特定的表示保持解耦。因此,我们提出了Gen2-VC,一种生成式视频压缩框架,它利用预训练的视频先验增强学习和传统编解码器的输出,同时保持其比特流和参考更新过程不变。在编解码器、VAE和生成骨干网络冻结的情况下,轻量级LoRA适配器通过单帧空间适配和多帧时间适配(结合视频先验)精化编解码器重建。使用Wan2.1-T2V-1.3B,Gen2-VC-DCVC-UF在LPIPS/DISTS指标上优于先前领先的编解码器,并且据我们所知,它是首个在低比特率下基于六个数据集的BD-rate平均值在PSNR和MS-SSIM上均超越VTM-23.0的生成式视频编解码器。与VTM-23.0相比,在匹配LPIPS和DISTS时,它分别平均降低比特率86.65%和94.24%。仅在DCVC-UF上训练的适配器无需重新训练即可提升DCVC-RT、ECM、VTM和HM的感知质量。

英文摘要:

Under stringent bitrate constraints, existing video codecs struggle to balance source fidelity and perceptual realism. Distortion-oriented codecs often oversmooth details, while generative codecs risk introducing content and structural deviations and rely on codec-specific designs. This motivates a question: Can existing codecs achieve a better distortion--perception trade-off through simple, reusable adaptation? Our insight is that native codec reconstructions provide a shared interface through which generative refinement is anchored to source content while remaining decoupled from codec-specific representations. We therefore propose Gen2-VC, a generative video compression framework that enhances the outputs of learned and conventional codecs with a pretrained video prior, leaving their bitstreams and reference update processes unchanged. With the codec, VAE, and generative backbone frozen, lightweight LoRA adapters refine codec reconstruction through single-frame spatial adaptation followed by multi-frame temporal adaptation with the video prior. Using Wan2.1-T2V-1.3B, Gen2-VC-DCVC-UF outperforms previous leading codecs in LPIPS/DISTS and, to our knowledge, is the first generative video codec to surpass VTM-23.0 in both PSNR and MS-SSIM at low bitrates, based on BD-rates averaged over six datasets. Compared to VTM-23.0, it reduces bitrate by an average of 86.65% and 94.24% at matched LPIPS and DISTS, respectively. Adapters trained only on DCVC-UF improve perceptual quality on DCVC-RT, ECM, VTM, and HM without retraining.

↑