arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13391cs.CV

上下文匹配蒸馏:用于自回归视频蒸馏的教师因果性

Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation

Hmrishav Bandyopadhyay, Xuanchi Ren, Zijian Huang, Jay Zhangjie Wu, Tianshi Cao, Ruilong Li, Bryan Chu, Sanja Fidler, Yi-Zhe Song, Zian Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出上下文匹配蒸馏(CMD)框架,解决现有视频蒸馏中教师监督与学生因果信息集不一致的问题,实现了自回归视频生成的最先进性能及对相机控制的更好依从性。

中文摘要 AI 辅助

交互式自回归视频生成既需要低延迟的推进,又需要精确的在线控制。少步蒸馏通过减少去噪步骤来加速生成,而在线控制则施加了因果约束:帧和块应依赖于生成过程中可用的历史和控制。然而,现有的视频分布匹配蒸馏(DMD)流水线通常使用对完整片段进行评分的双向教师来监督因果少步学生。因此,目标的评分可能依赖于学生生成时不可用的未来帧和控制,导致教师监督与学生的因果信息集不一致。我们引入上下文匹配蒸馏(Context-Matched Distillation,CMD),这是一种因果DMD框架,可使教师监督与每个目标生成时可用的信息保持一致。CMD用因果教师取代双向全片段评分,该教师在不访问未来帧或控制的情况下评估每个目标。同一因果教师初始化少步学生,在教师训练、学生蒸馏和推理过程中建立一致的因果表述。除了对齐时间信息边界外,前缀评分(Prefix Scoring)通过在产生目标的缓存学生生成前缀下评估每个目标,使监督与学生实际的推进上下文相匹配。前缀损坏(Prefix Corruption)通过在保留目标-上下文对齐的同时扰动训练早期产生的不可靠前缀,进一步稳定训练。凭借简单的因果表述,CMD自然扩展到逐帧和逐块生成、长视频蒸馏以及相机条件蒸馏。实验表明,CMD在短视频和长视频基准上的自回归方法中实现了最先进的聚合性能,同时大幅提高了对时变相机控制的依从性。

英文摘要

Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step distillation accelerates generation by reducing denoising steps, while online control imposes a causal constraint: frames and blocks should depend on history and controls available during generation. Existing video distribution matching distillation (DMD) pipelines, however, often supervise causal few-step students using bidirectional teachers that score complete clips. The score for a target can therefore depend on future frames and controls that were unavailable when the student generated it, misaligning teacher supervision with the student's causal information set. We introduce Context-Matched Distillation (CMD), a causal DMD framework that aligns teacher supervision with the information available when each target is generated. CMD replaces bidirectional full-clip scoring with a causal teacher that evaluates each target without access to future frames or controls. The same causal teacher initializes the few-step student, establishing a consistent causal formulation across teacher training, student distillation, and inference. Beyond aligning the temporal information boundary, Prefix Scoring matches supervision to the student's realized rollout context by evaluating each target under the cached student-generated prefix that produced it. Prefix Corruption further stabilizes training by perturbing unreliable prefixes produced early in training while preserving this target-context alignment. With a simple causal formulation, CMD naturally extends to frame-wise and chunk-wise generation, long video distillation, and camera-conditioned distillation. Experiments demonstrate state-of-the-art aggregate performance among autoregressive methods on both short- and long-video benchmarks, together with substantially improved adherence to time-varying camera controls.

发表机构

  • NVIDIA(英伟达)
  • University of Surrey(萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑