arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19437eess.IVcs.CVcs.MM

潜在组:通过掩码潜在生成建模实现极端比特率下的感知视频压缩

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

  • State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
  • State Key Discipline Laboratory of Wide Band-Gap Semiconductor Technology, School of Microelectronics, Xidian University, Xi'an, China(宽带隙半导体技术重点学科实验室,微电子学院,西安电子科技大学,西安,中国)

机构由 AI 辅助整理,请以论文原文为准。

Shaokang Wang, Jinchang Xu, Peidong Jia, Zhijian Hao, Siyuan Qian, Fei Zhao, Rui Ma, Xiaozhu Ju, Jian Tang, Xiaodong Xie, Shanghang Zhang, Huizhu Jia

AI总结:

针对极低比特率视频压缩未充分探索的问题,提出统一生成框架,通过潜在组策略、深度压缩模块和统一潜在去噪模块,在极低比特率下实现高感知质量,有丰富空间细节和强大时间一致性,达到最新水平。

AI中文摘要:

大多数现有视频压缩算法遵循变换和量化范式,优化失真和比特率之间的权衡。然而,极低比特率压缩仍是未充分探索的领域,严重受限编码资源下的感知质量优化未得到充分解决。本文提出一个统一生成框架,利用预训练的扩散Transformer先验在极低比特率下实现高感知质量。首先在因果分词器的潜在空间中引入灵活的潜在组策略,将潜在流明确分为帧内I潜在和帧间P潜在。深度压缩模块编码关键I潜在以最小开销保留感知锚点。基于这些锚点,基于DiT的统一潜在去噪模块细化帧内纹理并从噪声合成P潜在,以零额外比特率成本重建时间动态。大量实验表明该方法在极低比特率(如<0.005 bpp)下独特运行,实现具有丰富空间细节和强大时间一致性的最新感知保真度。代码将公开可用。

英文摘要:

Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitrate. However, extremely low-bitrate compression remains an underexplored frontier where perceptual quality optimization under severely constrained coding resources has not been adequately addressed. In this paper, we propose a unified generative framework that leverages pre-trained Diffusion Transformer (DiT) priors to achieve high perceptual quality at extremely low bitrates. We first introduce a flexible Group-of-Latents (GoL) strategy within the latent space of a causal tokenizer, explicitly partitioning the latent stream into intra $I$-latents and inter $P$-latents. The Deep Compression Module (I-DCM) then encodes key $I$-latents to preserve perceptual anchors with minimal overhead. Building upon these anchors, the DiT-based Unified Latent Denoising Module (U-LDM) refines intra-frame textures and synthesizes $P$-latents from noise, reconstructing temporal dynamics at zero additional bitrate cost. Extensive experiments demonstrate that our method uniquely operates in the extreme-low-bitrate regime (e.g., (<0.005) bpp), achieving state-of-the-art perceptual fidelity with rich spatial details and robust temporal consistency. The code will be made publicly available.

↑