发表机构
Tsinghua University; Kuaishou Technology; Sun Yat-sen University(清华大学; 快手科技; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到视频扩散模型的时间稀疏伪影问题,本文提出集中式隐式偏好优化(cIPO)框架,通过回退误差推导隐式偏好信号,将优化集中于高误差片段,有效提升了视频的真实性与时间连贯性。
AI 中文摘要
近期扩散式视频生成的偏好对齐技术,尤其是直接偏好优化(DPO)的进展,显著提升了视觉质量。但运动崩溃、对象闪烁、色彩过饱和等时间稀疏伪影仍是感知真实性的主要障碍。现有方法受两大局限制约:(1)偏好归因瓶颈,离线人工标注成本高且无法准确捕捉学习动态,在线奖励信号虽考虑回退但常不稳定且有偏差;(2)时间信用分配不当,均匀施加的监督无法精准针对伪影出现的短暂片段。为解决这些挑战,本文提出集中式隐式偏好优化(cIPO),这是一种针对视频扩散模型的后训练框架。cIPO直接从去噪过程推导隐式偏好信号:给定真实视频,模型添加前向噪声并通过迭代去噪重构,将原始视频视为偏好样本,重构结果视为非偏好样本。该公式无需人工标注或外部奖励模型即可捕捉推理时的误差。此外,原始视频与重构视频的帧级差异可揭示故障发生时机,cIPO利用这一点计算时间重构误差,并将优化集中于高误差片段,从而更精准地校正易故障区域。大量实验表明,cIPO在多个数据集上持续提升视频真实性和时间连贯性,凸显了带时间集中优化的隐式偏好的有效性与效率。
英文摘要
Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit preference with temporally concentrated optimization.
Commentsproject page: https://henglin-liu.github.io/cIPO_vis/