arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37850cs.CV

RelayVSR:大小模型协作实现高效真实世界视频超分辨率

RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution

  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Xijun Wang, Xin Li, Zirui Lang, Suhang Yao, Haoran Li, Zhibo Chen

AI总结:

RelayVSR提出稀疏生成中继机制,结合大型生成模型与轻量级VSR网络,通过视频感知参考优化提升真实世界视频超分辨率效率,在1080p下达到29.29 FPS。

AI中文摘要:

大型生成模型能够恢复真实世界视频超分辨率(VSR)中的真实细节,但使用它们处理整个视频在计算上代价高昂。在这项工作中,我们提出了RelayVSR,一个基于稀疏生成中继(Sparse Generative Relay)机制的流式VSR框架。大型生成模型为稀疏关键帧生成参考潜在表示,而轻量级VSR网络利用这些参考和低分辨率视频对每一帧进行超分辨率处理。轻量级VSR网络采用双记忆视频Transformer实现,跨帧重用关键帧信息并更新近期视频上下文,支持首关键帧条件化和具有有限前瞻的双端点条件化。然而,共享关键帧中的错误可能跨输出帧传播和累积,使得仅优化关键帧质量成为不充分的优化目标。我们通过视频感知参考优化(Video-Aware Reference Optimization,VARO)解决这一协作差距,该方法使用强化学习更新大型生成模型,具有两个奖励级别:系统级奖励评估由固定的轻量级VSR网络生成的视频,而参考级奖励评估解码后的关键帧质量。VARO相比直接联合训练提高了最终视频质量,其双级奖励优于仅使用系统级奖励。在单个NVIDIA A100 80GB上处理1080p视频时,具有15帧关键帧间隔的双端点RelayVSR达到29.29 FPS、13.82 GB峰值GPU内存和0.327秒首帧模型延迟,而FlashVSR-Tiny分别为7.80 FPS、24.447 GB和2.83秒。代码可在该https URL获取。

英文摘要:

Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at https://github.com/kopperx/RelayVSR.

补充信息

↑