FastVR:基于一步扩散的高效流式视频修复
FastVR: Efficient Streaming Video Restoration with One-Step Diffusion
浏览论文内容
中文总结 AI 辅助
FastVR提出一步扩散流式视频修复框架,结合轻量VAE与分块因果注意力,在H20 GPU上以11 FPS处理1080p视频,兼顾效率与质量,达到最先进性能。
中文摘要 AI 辅助
基于扩散的视频修复能够恢复真实细节,但其实际部署受到两个效率瓶颈的限制:昂贵的VAE编码和解码,以及扩散变换器(DiTs)中全自注意力的二次成本。本文提出FastVR,一种基于一步扩散模型的流式视频修复框架,在单个H20 GPU上以11 FPS处理1080p视频的同时,提供强大的修复质量和时间一致性。为提高推理效率,FastVR结合了轻量级VAE与分块因果注意力,大幅降低了计算成本。在训练过程中,它进一步采用速度一致性正则化和连续轨迹学习,以提升修复质量。大量实验表明,FastVR比所评估的扩散基线更高效,同时在合成和真实世界基准上取得了最先进的性能。我们希望这项工作能促进社区的进一步发展。
英文摘要
Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper presents FastVR, a streaming video restoration framework built on a one-step diffusion model, which delivers strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. To improve inference efficiency, FastVR combines a lightweight VAE with chunk-wise causal attention, which substantially reduces the computational cost. During training, it further adopts velocity consistency regularization and continuous trajectory learning, which improve restoration quality. Extensive experiments show that FastVR is more efficient than the evaluated diffusion baselines while achieving state-of-the-art performance on synthetic and real-world benchmarks. We hope that this work supports further progress in the community.
发表机构
- Alibaba Group(阿里巴巴集团)
- Xidian University(西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。