arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

增强自回归视频生成:基于表示对抗蒸馏

Enhancing Autoregressive Video Generation via Representation Adversarial Distillation

Fangyu Lin, Xingtong Ge, Lunjie Zhu, Yi Zhang, Zhening Liu, Tianhang Wang, Mengfei Li, Yumeng Zhang, Guanglu Song, Yu Liu, Jun Zhang

arXiv 2609.40037首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Vivix Group Limited; Zhejiang University(香港科技大学; Vivix集团有限公司; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Radian,一种表示空间对抗蒸馏框架,通过冻结视觉基础模型的特征对抗监督补充DMD,提升少步自回归视频生成的感知质量与稳定性。

AI 中文摘要

少步自回归视频生成能够实现高效的流式合成,但早期时间块中引入的错误会被作为上下文复用,并在后续的展开过程中传播,导致细节退化、结构漂移和运动不稳定。现有的分布匹配蒸馏(DMD)主要在扩散潜空间中对齐学生和教师分布,但对解码视频的感知质量缺乏直接监督。我们提出Radian,一种表示空间对抗蒸馏框架,通过在由冻结的视觉基础模型(VFM)定义的特征空间中,用真实数据对抗监督来补充在线策略DMD。在训练期间,Radian从自回归学生展开中稀疏地解码帧,提取多层级视觉表示,并应用轻量级判别器头来区分生成输出与真实视频帧。DMD目标将学生锚定到预训练教师,而表示空间对抗目标提供互补的感知和语义梯度,以促进高质量模式。这些附加组件在训练后被丢弃,生成器架构和推理时的去噪预算保持不变。在Wan2.1-1.3B上的实验涵盖四步分块、单步逐帧和分钟级自回归生成。我们的方法在四步生成下达到VBench总分0.8444和VideoAlign总分0.8033,并将VBench-Long从0.7805提升至0.8041,相比Rolling Forcing使用更少的去噪步骤。跨图像、视频和扩散表示的受控比较进一步表明,表示空间的选择会引发不同的对抗信号,且外部VFM梯度比源自扩散内部特征的对抗监督更能有效补充DMD。

英文摘要

Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, structural drift, and unstable motion. Existing distribution matching distillation (DMD) primarily aligns student and teacher distributions in diffusion latent space, but provides no direct supervision over the perceptual quality of decoded videos. We introduce Radian, a representation-space adversarial distillation framework that complements on-policy DMD with real-data adversarial supervision in the feature space defined by a frozen visual foundation model (VFM). During training, Radian sparsely decodes frames from autoregressive student rollouts, extracts multi-level visual representations, and applies lightweight discriminator heads to distinguish generated outputs from real video frames. The DMD objective anchors the student to the pretrained teacher, while the representation-space adversarial objective supplies complementary perceptual and semantic gradients that promote high-quality modes. These additional components are discarded after training, leaving the generator architecture and inference-time denoising budget unchanged. Experiments on Wan2.1-1.3B cover four-step chunk-wise, one-step frame-wise, and minute-long autoregressive generation. Our method achieves a VBench Total of 0.8444 and a VideoAlign Total of 0.8033 under four-step generation, and improves VBench-Long from 0.7805 to 0.8041 over Rolling Forcing while using fewer denoising steps. Controlled comparisons across image, video, and diffusion representations further indicate that the choice of representation spaces induces distinct adversarial signals, and external VFM gradients complement DMD more effectively than adversarial supervision derived from diffusion-internal features.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑