LongLive-Plug:视频生成的一次性蒸馏框架
LongLive-Plug: Once-for-All Distillation for Video Generation
浏览论文内容
中文总结 AI 辅助
提出LongLive-Plug,一种视频生成的一次性蒸馏框架,将可复用能力(CFG、少步采样、长上下文纠错)以LoRA形式学习,免训练部署至54个下游模型,避免逐模型重复蒸馏。
中文摘要 AI 辅助
视频扩散模型正日益发展为面向多种下游任务的专用模型,这一发展过程通常包含一个蒸馏阶段,例如用于加速采样或改进长视频生成。这一阶段通常需要对每个专用模型重复进行。我们提出了LongLive-Plug,一个一次性蒸馏框架,该框架在基础模型上以LoRA形式学习可复用能力,实现免训练、即插即用地部署到兼容的下游模型。这些能力包括单遍无分类器引导、少步采样以及自回归生成的长上下文错误修正。即使下游模型添加了条件分支、扩展了输出通道,这些适配器仍然保持可复用性。尽管在固定的引导尺度下进行训练,我们专门的CFG LoRA通过其推理权重提供文本引导控制。将其与少步LoRA结合,可以在下游任务上同时保留少步生成能力和CFG可控性。我们验证了在三个骨干家族和八个任务类别(包括世界建模、机器人、编辑和多模态生成)中的54个下游模型上的免训练部署。该方法可能支持额外的兼容模型。因此,每种能力可以在每个骨干家族上蒸馏一次,而无需针对每个目标进行重新训练。
英文摘要
Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.