arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReCaVSR:利用回收潜变量与学习型缓存路由的单步流式扩散视频超分辨率

ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing

Xijun Wang, Xin Li, Suhang Yao, Zirui Lang, Bingchen Li, Zhibo Chen

arXiv 2609.37831首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReCaVSR提出基于Wan2.2的单步流式扩散视频超分辨率框架,通过回收潜变量与学习型缓存路由减少延迟,结合多范围判别器和LR条件化解码器,在保持生成质量的同时实现高效流式处理。

AI 中文摘要

实时扩散视频超分辨率(VSR)在在线流媒体中需求旺盛,但严格的延迟要求往往损害生成保真度。我们提出ReCaVSR,一个基于Wan2.2的单步流式VSR框架,其构建于两个观察之上:回收的SR潜变量保留了局部时间上下文,减少了对完整历史键值(KV)缓存的需求;以及各个Transformer层受益于不同的时间范围。ReCaVSR结合了三种互补设计:(i)带回收SR潜变量的逐层缓存路由:每个DiT层在缓存预算下学习其KV缓存时间范围,并导出静态推理调度,而回收的SR潜变量通过将每个新块以模型自身先前预测为条件来传播局部上下文。(ii)多范围查询(MSQ)判别器:一种组合判别器,结合全局、空间窗口和时间管反馈,以实现整体真实性、局部纹理生成和时间稳定性。(iii)LR条件化的FlashDecoder适配:一种VAE解码器,结合LR观测以实现高效潜变量解码。ReCaVSR无需迭代采样或完整历史KV缓存物化即可实现流式VSR。在合成和真实世界VSR基准上的实验显示,与代表性VSR基线相比,具有更好的感知质量、时间一致性和流式效率。在单个NVIDIA A100-80GB上以$1080\ imes1920$输出分辨率,ReCaVSR达到21.20 FPS,峰值分配GPU内存15.16 GB,比FlashVSR Tiny快2.72倍,同时峰值分配内存减少38.0%。代码可在以下https URL获取。

英文摘要

Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At $1080{\times}1920$ output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72$\times$ faster while using 38.0\% less peak allocated memory than FlashVSR Tiny. The code is available at https://github.com/kopperx/ReCaVSR.

CommentsThe code is available at https://github.com/kopperx/ReCaVSR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑