发表机构
Adobe Research; University of Rochester(Adobe研究院; 罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出EditStream框架,整合多类视频生成编辑任务,通过两阶段蒸馏方法优化自回归模型,实现高质量交互式视频创作,填补了扩散模型与创意工作流的衔接空白。
AI 中文摘要
交互式视频生成与编辑在创意设计领域的重要性日益提升。本研究提出EditStream——一个用于交互式视频生成与编辑的统一框架。EditStream通过灵活的任务特定条件控制,在单个基于DiT的模型中整合了多种视频创建与操作任务,并将其转化为快速、少步长的自回归模型以实现高效流处理。它支持文本生成视频、图像生成视频、视频转视频、编辑传播、参考引导视频编辑以及相机姿态变更,可在单一系统内实现对视频生成、转换与编辑的灵活控制。为使该统一模型适用于交互式场景,本研究开发了两阶段蒸馏方法,将速度矩匹配(Velocity Moment Matching, VMM)与自回归展开相结合。VMM在学生模型到达的中间状态匹配条件速度矩,以保留生成质量与运动;自回归展开则让学生模型接触自身的自回归预测,以提升时间稳定性。二者共同缓解了少步长自回归视频生成中的常见挑战,包括过饱和、运动退化、时间不稳定及训练复杂度高的问题。EditStream提供了一种实用且可扩展的解决方案,弥合了高质量扩散型视频模型与交互式创意工作流之间的差距。
英文摘要
Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible task-specific conditioning, and further transforms it into a fast, few-step autoregressive model for efficient streaming. It supports Text-to-Video, Image-to-Video, Video-to-Video, Editing Propagation, Reference-guided Video Editing, and Camera Pose Change, enabling flexible control over video generation, transformation, and editing within one system. To make the unified model practical for interactive use, we develop a two-stage distillation approach that combines Velocity Moment Matching (VMM) with autoregressive unrolling. VMM matches conditional velocity moments at student-reached intermediate states to preserve generation quality and motion, while unrolling exposes the student to its own autoregressive predictions to improve temporal stability. Together, they alleviate common challenges in few-step autoregressive video generation, including over-saturation, degraded motion, temporal instability, and complex training. EditStream provides a practical and scalable solution that bridges high-quality diffusion-based video models with interactive creative workflows.
Comments25 pages, 12 figures, Project page: https://real-time-video-research.github.io/editstream/