arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21268cs.CV

Edit-VAR:驯服视觉自回归模型以实现精确视频编辑

Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

  • Sun Yat-sen University(中山大学)
  • South China University of Technology(华南理工大学)
  • Tsinghua University(清华大学)
  • Shandong University(山东大学)
  • Nankai University(南开大学)
  • Hunan University(湖南大学)

机构由 AI 辅助整理,请以论文原文为准。

Chongbo Zhao, Jiangming Wang, Xilai Wang, Xinyu Wang, Jingyi Tang, Chunjie Hao, Pengjie Song, Yue Ma

AI总结:

Edit-VAR提出首个免训练免反演的视觉自回归视频编辑框架,通过概率引导令牌替换和尺度解耦生成,实现高保真、源保留且高效的视频编辑。

AI中文摘要:

文本引导的视频编辑在修改目标内容的同时,保持未编辑区域的外观和时间一致性。基于训练的方法提供了强大的控制力,但需要大量的数据和计算资源。免训练方法分为免反演和基于反演两种范式。免反演方法避免了轨迹恢复,但其源保留引导可能限制编辑强度,并使语义变化不完整。基于反演的方法在再生成前恢复潜在轨迹,其中近似误差可能累积,导致源内容漂移和时间不一致。我们提出了Edit-VAR,这是首个利用预训练视觉自回归视频模型进行文本引导视频编辑的免训练且免反演框架。Edit-VAR直接将源视频编码为多尺度离散令牌,并执行概率引导的条件令牌替换以实现源保留。注意力引导的逐令牌和尺度感知调制选择性地放宽编辑相关位置和生成阶段上的源约束。尺度解耦生成(实现为后期尺度约束释放)重新生成运动一致的细节并减少纹理碎片化。残差引导的令牌剪枝进一步利用最后两个高分辨率尺度上的冗余来降低推理成本。大量实验和一项盲用户研究表明,Edit-VAR在编辑保真度、源保留、时间一致性和推理效率方面总体上优于现有的免训练视频编辑方法。

英文摘要:

Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.

补充信息

↑