arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OSVE:使用单步扩散模型进行单步视频编辑

OSVE: One Step Video Editing with One Step Diffusion Models

Habin Lim, Gyeong-Moon Park

arXiv 2607.19895首次发表:更新:

发表机构

Korea University(韩国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散模型文本引导视频编辑慢的问题,提出OSVE框架。通过训练可学习编码器预测初始噪声,引入UFE技术及滑动窗口策略保持一致性,实现高质量视频编辑,速度比现有方法快约155 - 171倍。

AI 中文摘要

使用扩散模型进行文本引导的视频编辑速度极慢,受多步采样和反演成本的阻碍。我们提出了OSVE,这是第一个成功将单步文本到图像(T2I)模型用于高质量视频编辑的框架,解决了反演、可编辑性和时间一致性的核心挑战。为绕过缓慢的迭代反演,我们训练了一个可学习的编码器,在单个前向传播中预测每一帧的初始噪声。该编码器在精心策划的结构对齐图像对数据集上,用一种新颖的结构感知编辑(SAE)损失进行训练,使其在编辑过程中保留源视频的几何结构。对于时间连贯性,我们引入了统一帧编辑(UFE)技术,在单个生成步骤中连接帧潜在表示以促进跨帧注意力。此外,对于长视频,采用带有锚帧的滑动窗口策略保持全局一致性。我们的大量实验表明,OSVE实现的编辑质量与最先进的多步方法相当或更优,同时运行速度快约155 - 171倍。这一突破为实用的实时视频编辑应用铺平了道路。

英文摘要

Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion. We present OSVE, the first framework to successfully adapt one-step Text-to-Image (T2I) models for high-quality video editing, addressing the core challenges of inversion, editability, and temporal consistency. To bypass slow iterative inversion, we train a learnable encoder that predicts the initial noise for each frame in a single forward pass. This encoder is trained with a novel Structure-Aware Editing (SAE) loss on a curated dataset of structurally-aligned image pairs, teaching it to preserve the source video's geometry during edits. For temporal coherence, we introduce Unified-Frame Editing (UFE), a technique that concatenates frame latents to facilitate cross-frame attention in a single generation step. Furthermore, for long videos, a sliding-window strategy with an anchor frame maintains global consistency. Our extensive experiments demonstrate that OSVE achieves editing quality comparable or superior to state-of-the-art multi-step methods, while operating approximately 155--171 times faster. This breakthrough paves the way for practical, real-time video editing applications. Code is available at https://github.com/KU-VGI/OSVE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑