arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TPD:用于文本到视频扩散模型的时间先验解耦

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models

Taewon Kang, Matthias Zwicker

arXiv 2607.26706首次发表:更新:

AI 中文总结

该研究针对文本到视频扩散模型的时间先验抑制问题,提出无需训练的TPD框架,通过帧选择性下界约束恢复被抑制的后期片段信号,提升后期概念实现效果且与骨干无关。

AI 中文摘要

文本到视频扩散模型可根据自然语言生成时间连贯的内容,但当提示描述的早期场景持续存在,同时新事件在其之上出现时——例如“一座高大的沙堡矗立在海滩上,海浪冲来将其卷走”——生成结果常无法在对应帧中实现后期片段的事件。我们将此失败现象识别为时间先验抑制(Temporal Prior Suppression, TPS):早期片段的主导先验捕获了时间轴上的交叉注意力轨迹,并抑制了实现后期片段所需的引导信号,这是现有引导机制未建模的竞争倾向。我们提出时间先验解耦(Temporal Prior Decoupling, TPD),这是一种无需训练的框架,可在扩散采样期间恢复被抑制的后期片段信号。TPD通过仅以早期片段为条件构建时间反事实,并将完整提示与反事实轨迹之间的差异定义为被抑制的信号方向。与先前的减法投影方法去除该方向不同,TPD通过在扩散时间步长和视频帧上联合求解的帧选择性下界约束来恢复该方向,在不破坏早期片段连贯性的情况下实现后期帧中的被抑制事件:先前工作采用上界可行性去除不需要的语义,而TPD采用下界可行性保证被抑制信号的贡献。TPD完全在标准扩散采样内运行,无需重新训练,且仅在无分类器引导空间中定义,因此本质上与骨干无关。实验表明,TPD在保留时间连贯性和视觉保真度的同时显著提升了后期概念的实现,且这种针对性抑制在不同的文本到视频骨干中均存在。

英文摘要

Text-to-video diffusion models generate temporally coherent content from natural language, yet when a prompt describes an early scene that persists while a new event emerges on top of it---such as "a tall sandcastle standing on a beach where a wave rushes in and washes it away"---generation frequently fails to realize the late-segment event in the corresponding frames. We identify this failure as Temporal Prior Suppression (TPS): the dominant prior of the early segment captures the cross-attention trajectory across the temporal axis and suppresses the guidance signal needed for late-segment realization, a competing tendency existing guidance mechanisms do not model. We introduce Temporal Prior Decoupling (TPD), a training-free framework that restores suppressed late-segment signals during diffusion sampling. TPD constructs a temporal counterfactual by conditioning on the early segment alone, and defines the discrepancy between the full-prompt and counterfactual trajectories as a suppressed signal direction. Rather than removing this direction as in prior subtractive projection methods, TPD restores it through a frame-selective lower-bound constraint resolved jointly over diffusion timestep and video frame, realizing the suppressed event in the late frames without disrupting early-segment coherence: where prior work enforces upper-bound feasibility to remove unwanted semantics, TPD enforces lower-bound feasibility to guarantee suppressed-signal contribution. TPD runs entirely within standard diffusion sampling without retraining, and is defined purely in classifier-free guidance space, making it backbone-agnostic by construction. Experiments show that TPD significantly improves late-concept realization while preserving temporal coherence and visual fidelity, and that the targeted suppression recurs across distinct text-to-video backbones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑