arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04891eess.SP

GVC-RT:面向超低码率下的实时生成视频压缩

GVC-RT: Towards Real-Time Generative Video Compression at Ultra-Low Bitrates

Tianjian Dang, Sixian Wang, Lei Luo, Guo Lu, Jincheng Dai

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有生成视频编解码器实时性不足的问题,提出GVC-RT,通过重新设计生成式隐编码框架,在超低码率下实现了优于SOTA模型的压缩性能与实时编码解码速度。

中文摘要 AI 辅助

近期的生成视频编解码器(Generative Video Codecs, GVC)通过压缩生成式分词器(tokenizer)生成的token,在超低码率(<0.02比特每像素)下实现了令人印象深刻的重建保真度。然而,现有GVC通常需要大量计算时间和模型复杂度,这阻碍了它们在计算受限设备和实时应用中的部署。为弥合这一差距,本文系统识别了计算瓶颈并提出GVC-RT,其重新设计了生成式隐编码框架,在不牺牲压缩性能的前提下实现了实时视频编码。具体而言,GVC-RT基于预训练的无查找量化(Lookup-Free Quantization, LFQ)分词器构建,采用非对称架构直接学习匹配LFQ隐分布,且仅在训练阶段通过正则化损失项强制生成空间对齐。这种方式绕过了繁重的分词操作,并在推理阶段完全移除了复杂的特征对齐过程。此外,本文进一步引入轻量化解分词器(de-tokenizer)架构以解决解码阶段的最终延迟瓶颈。实验结果表明,GVC-RT在DISTS和LPIPS指标上分别实现了12.4%和48.8%的平均BD-rate节省,优于此前的SOTA模型GLC-Video,同时在1080p视频上实现了123.1帧每秒的编码速度和55.1帧每秒的解码速度。代码位于此https URL。

英文摘要

Recent generative video codecs (GVCs) have achieved impressive reconstruction fidelity at ultra-low bitrates (< 0.02 bits per pixel) by compressing the tokens from generative tokenizers. However, existing GVCs generally require considerable computation time and model complexity, which hinder their deployment on compute-limited devices and in real-time applications. To bridge this gap, we systematically identify the computational bottlenecks and propose GVC-RT, which redesigns the generative latent coding framework to realize real-time video coding without sacrificing compression performance. Specifically, built on a pretrained lookup-free quantization (LFQ) tokenizer, GVC-RT adopts an asymmetric architecture that directly learns to match the LFQ latent distribution, while generative-space alignment is enforced via a regularization loss term only during training. In this manner, we bypass heavy tokenization and entirely remove the complex feature-alignment process at inference time. Moreover, we further introduce a lightweight de-tokenizer architecture to resolve the final latency bottleneck during decoding. Experimental results demonstrate that GVC-RT outperforms the previous SOTA model, GLC-Video, with average BD-rate savings of 12.4% and 48.8% in terms of DISTS and LPIPS, while achieving encoding/decoding speeds of 123.1/55.1 fps for 1080p video. The code is at https://github.com/semcomm/GVC-RT.

补充信息

↑