arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于文本到图像潜在扩散模型与多模态参考的零样本视频恢复与增强

Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

Cong Cao, Huanjing Yue, Xin Liu, Jingyu Yang

arXiv 2608.26476首次发表:更新:

发表机构

School of Electrical and Information Engineering, Tianjin University; Lappeenranta-Lahti University of Technology LUT(天津大学电气自动化与信息工程学院; 拉彭兰塔-拉赫蒂理工大学(LUT))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对文本到图像潜在扩散模型用于视频恢复时出现的时间闪烁问题,提出结合多模态参考的零样本视频恢复增强框架,通过双提示调优反演采样等技术提升效率与时间一致性,实验验证其优越性。

AI 中文摘要

采用文本到图像潜在扩散模型的零样本图像恢复方法在无需训练的通用图像恢复任务中取得了巨大成功,但将其应用于视频恢复会产生严重的时间闪烁问题。本文提出一种用于零样本视频恢复与增强的新型框架,该框架采用文本到图像潜在扩散模型与多模态参考。通过所提出的双提示调优反演与采样,推理时间可缩短至原来的近1/3,性能和时间一致性也得到显著提升。通过采用所提出的感知纹理的视频令牌合并,可进一步利用帧间的时间相关性以改善时间一致性。我们还提出了参考自注意力与参考令牌合并以支持图像参考。实验结果表明,所提方法在恢复和增强时间一致的视频方面具有优越性。

英文摘要

Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without training. However, applying them to video restoration will result in severe temporal flickering. In this paper, we propose a novel framework for zero-shot video restoration and enhancement which uses a text-to-image latent diffusion model and multi-modal references. Through the proposed dual prompt tuning inversion and sampling, the inference time can be reduced to nearly 1/3 of the original. The performance and temporal consistency can be also significantly stregthened. By using the proposed texture-aware video token merging, the temporal correlation between frames can be further utilized to improve the temporal consistency. We futher propose the referenced self-attention and referenced token merging to support image reference. Experimental results demonstrate the superiority of the proposed method in restoring and enhancing temporally consistent videos.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑