arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiffVC-ONE:基于扩散模型的生成式视频压缩,采用单步视频扩散Transformer

DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer

Wenzhuo Ma, Zhenzhong Chen

arXiv 2608.20515首次发表:更新:

发表机构

school of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感信息工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DiffVC-ONE是基于扩散模型的生成式视频压缩框架,采用单步视频扩散Transformer,通过三类组件实现低推理成本下的高感知质量与时间一致性,在多基准上达最优性能。

AI 中文摘要

生成式视频压缩可在低码率下恢复丰富的视觉细节,但同时实现高时间一致性和低推理成本仍具挑战性。为解决该问题,本文提出DiffVC-ONE,一种基于扩散模型的生成式视频压缩框架,构建于单步视频扩散Transformer(Video Diffusion Transformer)之上。首先,引入统一单向潜在压缩器(Unified Unidirectional Latent Compressor),使用共享模型高效且均匀地压缩紧凑的潜在切片;随后,开发基于视频DiT的单步扩散增强器(Video DiT-based One-Step Diffusion Enhancer),将重构后的潜在切片作为内容锚点,对整组图像执行单步时空感知增强;最后,混合条件生成器(Hybrid Condition Generator)从重构内容和量化信息中提取结构、强度及语义条件,这些条件在单步扩散增强过程中保留忠实区域、控制生成增强的程度并补充内容感知的感知细节。在多个标准基准上开展的大量实验表明,DiffVC-ONE以低推理成本实现了最先进的感知质量和时间一致性。

英文摘要

Generative video compression can recover rich visual details at low bitrates, but simultaneously achieving high temporal consistency and low inference cost remains challenging. To address this issue, we propose DiffVC-ONE, a diffusion-based generative video compression framework built on a one-step Video Diffusion Transformer. First, we introduce a Unified Unidirectional Latent Compressor that uses a shared model to efficiently and uniformly compress compact latent slices. We then develop a Video DiT-based One-Step Diffusion Enhancer that uses the reconstructed latent slices as content anchors and performs single-step spatio-temporal perceptual enhancement over an entire group of pictures. Finally, a Hybrid Condition Generator extracts structural, strength, and semantic conditions from the reconstructed content and quantization information. These conditions preserve faithful regions, control the degree of generative enhancement, and supplement content-aware perceptual details during one-step diffusion enhancement. Extensive experiments on multiple standard benchmarks demonstrate that DiffVC-ONE achieves state-of-the-art perceptual quality and temporal consistency with low inference cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑