VoRTeC:驯服基础流模型实现单步实时视频压缩
VoRTeC: Taming Foundation Flow for One-step Real time Video Compression
另 1 家 · 查看机构详情
- Shenzhen International Graduate School(深圳国际研究生院)
- Tsinghua University(清华大学)
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对超低码率视频压缩的模糊伪影、解码延迟与时间一致性问题,提出基于Wan2.1的VoRTeC框架,实现单步解码,比特消耗降58%,解码速度提升3至197倍。
中文摘要 AI 辅助
超低码率视频压缩仍面临关键挑战:传统神经视频压缩不可避免引入模糊伪影,而基于扩散的生成式视频压缩存在解码延迟过高、时间一致性差的问题。为解决这些问题,我们提出VoRTeC,这是一个基于基础流模型(Wan2.1)构建的视频压缩框架。通过紧凑编码视频潜表征、预测压缩表征沿流轨迹的位置,以及整合多尺度先验,VoRTeC使压缩器能有效利用生成式视频流先验。无需访问流匹配网络的参数或梯度,我们的框架实现单步解码,且重构具有高感知保真度。同时,我们通过尾帧复用和先验缓存维持帧组间的一致性。大量实验表明,与现有基于扩散的方法相比,我们的方法减少了58%的比特消耗,解码速度提升了3至197倍:VoRTeC在720p分辨率下实现13 FPS的解码速度,在480p分辨率下实现32 FPS的解码速度。
英文摘要
Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose $\mathtt{VoRTeC}$, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video representations, predicting the positions of compressed representations along flow trajectories, and integrating multi-scale priors, $\mathtt{VoRTeC}$ enables the compressor to harness generative video flow priors effectively. Without accessing the parameters or gradients of flow matching networks, our framework achieves one-step decoding and reconstructions with high perceptual fidelity. Meanwhile, we maintain consistency across frame groups via tail-frame reuse and prior caching. Extensive experiments demonstrate that our method reduces bit consumption by 58\% compared to prior diffusion-based approaches, with decoding speed boosted by 3 to 197 times: $\mathtt{VoRTeC}$ achieves a decoding speed of 13 FPS at 720p and 32 FPS at 480p.