arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21529cs.CVcs.AI

ElasticTTT:用于视频编辑的保留先验的测试时调优

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对预训练扩散模型的测试时调优中存在的问题,提出ElasticTTT框架,通过目标分布正则化、对比条件采样和异步噪声调度等方法,成功保留基础模型生成先验,在一次性视频编辑中达领先性能。

中文摘要 AI 辅助

预训练扩散模型上的测试时调优(TTT)已成为视频编辑的强大范式。然而,生成模型的分布映射性质与标准TTT的单点优化之间存在根本不匹配。本文证明这种不匹配会引发“先验崩溃”,即模型丢弃文本条件和空间潜在信息,使生成退化为源视频,或混淆不同区域的特征。为解决此问题,我们提出了ElasticTTT,这是一个保留先验生成分布并恢复生成弹性的新框架。具体而言,我们提出了目标分布正则化以防止尖锐的记忆最小值,对比条件采样以引导推理远离源偏差,以及异步噪声调度以保留未编辑区域。广泛的评估表明,ElasticTTT成功保留了基础模型的生成先验,在一次性视频编辑中实现了领先性能。

英文摘要

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.

发表机构

  • College of AI, Tsinghua University(清华大学人工智能学院)
  • Beijing Academy of Artificial Intelligence(北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑