发表机构
Information Technologies Institute, CERTH; University of Amsterdam(信息技术研究所,希腊研究与技术中心; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对自回归视频扩散模型误差积累问题,提出TANGO方法,通过噪声引导优化避免模型陷入终点,在VBench上有显著改进,降低了弗雷歇视频距离。
AI 中文摘要
自回归视频扩散模型通过消除对未来帧的条件依赖,能够生成任意长的视频,显著提高了计算效率。然而,由于去噪序列逐渐偏离训练期间看到的条件分布,它们会随着时间积累误差。最近的进展试图通过将每个生成的帧锚定到真实帧的学习流形上来减少这种误差。但即使所有生成的单个帧都靠近真实流形,模型仍缺乏足够知识来继续某些轨迹而不离开它,从而到达终点。为防止模型陷入终点,我们从这样的假设出发:对于建模良好的未来轨迹,预测噪声的分布应与前向噪声过程的分布相匹配。为在测试时强制实施这种先验,我们引入了通过噪声引导优化避免终点(TANGO),它将扩散模型用作自身输出的评判器,通过向前预测一步并要求各向同性高斯噪声预测。我们利用与预期噪声分布的偏差来搜索不会导致终点的替代轨迹。我们的方法在VBench上比现有技术实现了3.1%的绝对改进,同时在15秒视频上平均将弗雷歇视频距离降低了28.3%。我们的代码可在这个https URL上获取。
英文摘要
Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thus greatly improving computational efficiency. Yet, they suffer from error accumulation over time, as the denoised sequence gradually drifts away from the conditioning distribution seen during training. Recent advances attempt to reduce this error by anchoring each generated frame to the learned manifold of real ones. However, even when all generated individual frames lie close to the real manifold, there are trajectories which the model lacks sufficient knowledge to continue without exiting it, thus reaching a terminal point. To prevent the model from being trapped in terminal points, we start from the hypothesis that for well-modeled future trajectories the distribution of the predicted noise should match the one of the forward noising process. To enforce such a prior at test time, we introduce Terminal points Avoidance through Noise Guided Optimization (TANGO), which uses the diffusion model as a critic of its own outputs, by predicting one step forward and requiring an isotropic Gaussian noise prediction. We use the deviation from this expected noise distribution to search for an alternative trajectory that does not lead to a terminal point. Our approach achieves a $3.1\%$ absolute improvement on VBench over state-of-the-art, while reducing Fréchet Video Distance by $28.3\%$ on average across $15$s videos. Our code is available on https://mever-team.github.io/tango.
CommentsECCV2026