潜在组:通过掩码潜在生成建模实现极端比特率下的感知视频压缩
Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling
- State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
- State Key Discipline Laboratory of Wide Band-Gap Semiconductor Technology, School of Microelectronics, Xidian University, Xi'an, China(宽带隙半导体技术重点学科实验室,微电子学院,西安电子科技大学,西安,中国)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对极低比特率视频压缩未充分探索的问题,提出统一生成框架,通过潜在组策略、深度压缩模块和统一潜在去噪模块,在极低比特率下实现高感知质量,有丰富空间细节和强大时间一致性,达到最新水平。
AI中文摘要:
大多数现有视频压缩算法遵循变换和量化范式,优化失真和比特率之间的权衡。然而,极低比特率压缩仍是未充分探索的领域,严重受限编码资源下的感知质量优化未得到充分解决。本文提出一个统一生成框架,利用预训练的扩散Transformer先验在极低比特率下实现高感知质量。首先在因果分词器的潜在空间中引入灵活的潜在组策略,将潜在流明确分为帧内I潜在和帧间P潜在。深度压缩模块编码关键I潜在以最小开销保留感知锚点。基于这些锚点,基于DiT的统一潜在去噪模块细化帧内纹理并从噪声合成P潜在,以零额外比特率成本重建时间动态。大量实验表明该方法在极低比特率(如<0.005 bpp)下独特运行,实现具有丰富空间细节和强大时间一致性的最新感知保真度。代码将公开可用。
英文摘要:
Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitrate. However, extremely low-bitrate compression remains an underexplored frontier where perceptual quality optimization under severely constrained coding resources has not been adequately addressed. In this paper, we propose a unified generative framework that leverages pre-trained Diffusion Transformer (DiT) priors to achieve high perceptual quality at extremely low bitrates. We first introduce a flexible Group-of-Latents (GoL) strategy within the latent space of a causal tokenizer, explicitly partitioning the latent stream into intra $I$-latents and inter $P$-latents. The Deep Compression Module (I-DCM) then encodes key $I$-latents to preserve perceptual anchors with minimal overhead. Building upon these anchors, the DiT-based Unified Latent Denoising Module (U-LDM) refines intra-frame textures and synthesizes $P$-latents from noise, reconstructing temporal dynamics at zero additional bitrate cost. Extensive experiments demonstrate that our method uniquely operates in the extreme-low-bitrate regime (e.g., (<0.005) bpp), achieving state-of-the-art perceptual fidelity with rich spatial details and robust temporal consistency. The code will be made publicly available.