发表机构
Communication University of China; University of Science and Technology of China; Microsoft Research Asia; Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国传媒大学; 中国科学技术大学; 微软亚洲研究院; 中国科学院数学与系统科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视频压缩探索专门设计的扩散模型,提出GenVC,通过全局到局部层次结构恢复时空细节,用自适应分数蒸馏加速推理,在超低比特率下实现高质量重建,参数少且解码速度快。
AI 中文摘要
扩散模型为超低比特率视频压缩提供了强大的生成能力。现有基于扩散的视频编解码器采用最初为文本条件生成开发的基础模型,而专门为压缩设计和训练的扩散模型尚未被探索。为填补这一空白,我们引入了基于从零开始训练的视频扩散模型的生成式视频编解码器(GenVC)。我们通过全局到局部层次结构在像素空间中直接实现该模型,以恢复精细的时空细节。为加速推理,我们使用分布匹配蒸馏(DMD)将多步模型蒸馏为一步。然而,直接应用DMD会导致运动停滞的重建。我们将此归因于教师侧指导失败。为打破由此产生的反馈循环,我们提出了自适应分数蒸馏,根据与真实方向的对齐来控制DMD更新。实验结果表明,GenVC在超低比特率下实现了一流的感知质量。与之前继承数十亿规模预训练主干的编解码器不同,我们的扩散模型只有478.0M参数,在A100 GPU上以15.1 fps的速度单步解码1080p视频。
英文摘要
Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base models originally developed for text-conditioned generation, whereas diffusion models designed and trained specifically for compression remain unexplored. To fill this gap, we introduce our Generative Video Codec (GenVC), built on a video diffusion model trained from scratch for compression. To our knowledge, this is the first compression-oriented video diffusion model. We realize this model directly in pixel space with a global-to-local hierarchy that recovers fine spatio-temporal details, enabling high-quality generative reconstruction from compressed representations. To accelerate inference, we distill the multi-step model into one step using distribution matching distillation (DMD). Applying DMD directly, however, drives the student toward motion-stalled reconstructions. We trace this to a teacher-side guidance failure: once student-induced perturbations leave the frozen teacher's training region, its guidance can become misleading, causing DMD updates to reinforce rather than correct the student drift. To break the resulting feedback loop, we propose Adaptive Score Distillation, which gates DMD updates according to their alignment with the ground-truth direction, enabling high-quality reconstruction with coherent motion. Experimental results show that GenVC achieves state-of-the-art perceptual quality at ultra-low bitrates, with average bitrate savings of 62.5% at matched LPIPS and 71.3% at matched FID over GLVC. Unlike prior codecs that inherit billion-scale pretrained backbones, our diffusion model has only 478.0M parameters and decodes 1080p video in a single step at 15.1 fps on an A100 GPU.