HPSD:用于文本-图像转视频扩散模型的混合策略自蒸馏
HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models
浏览论文内容
中文总结 AI 辅助
本研究提出HPSD混合策略自蒸馏框架,通过让同一TI2V模型在不同条件下分别充当教师与学生,解决现有自蒸馏方法的缺陷,显著提升T2V及TI2V生成性能,强化模型基础生成能力。
中文摘要 AI 辅助
文本-图像转视频(TI2V)模型是一种新兴的统一架构,单个模型可同时支持文本转视频(T2V)和图像转视频(I2V)生成。在获得高质量首帧或详细文本提示的情况下,TI2V模型的视觉质量远优于其T2V模式,由此引出一个自然问题:能否将这类特权条件下激发的能力内化到模型自身的基础生成能力中?实现这一目标的常用方法是模型自蒸馏。然而,最直接的解决方案——监督微调采用离策略策略:其监督仅来自教师模型生成的固定离线分布端点,而非学生模型访问的状态,缺乏针对不断演变策略的精确修正。近期的在线策略蒸馏方法则存在条件-状态不匹配问题,监督指向给定的首帧,而非学生模型的实际内容,导致修正出现偏差。为实现既能吸收教师模型的特权先验,又能保留精确策略修正的自蒸馏,本研究提出混合策略自蒸馏(Hybrid-Policy Self-Distillation,HPSD),这是一种新型自蒸馏框架:单个TI2V模型在不同条件下分别充当教师和学生,教师在带有高质量首帧和增强提示的TI2V模式下运行,学生在仅使用原始提示的基础T2V模式下运行。具体而言,学生模型继承教师模型的离策略轨迹点作为锚点,针对自身策略进行局部优化,最终在这些自主生成的滚动输出上接收速度级别的监督。大量实验表明,HPSD可显著提升T2V性能,同时也能带来显著的TI2V增益,有效增强模型的基础生成能力。
英文摘要
Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given a high-quality first frame or a detailed textual prompt, TI2V models unlock substantially better visual quality than their T2V mode, raising a natural question: can the capability elicited by such privileged conditions be internalized into the model's own base generation ability? A common approach toward this goal is model self-distillation. However, the most straightforward solution, supervised fine-tuning, follows an off-policy strategy: its supervision is confined to teacher-generated endpoints from a fixed offline distribution rather than student-visited states, lacking precise correction tailored to the evolving policy. Recent on-policy distillation methods instead suffer from condition-state mismatch, where supervision is steered toward the given first frame instead of the student's actual content, misleading the correction. To achieve self-distillation that absorbs the teacher's privileged prior while retaining precise policy correction, in this work, we propose Hybrid-Policy Self-Distillation (HPSD), a novel self-distillation framework where a single TI2V model acts as both teacher and student under different conditions: the teacher operates in TI2V mode with a high-quality first frame and an enhanced prompt, while the student runs in the base T2V mode with only the vanilla prompt. Specifically, the student inherits off-policy teacher trajectory points as anchors, locally refines them toward its own policy, and finally receives velocity-level supervision on these self-generated roll-outs. Extensive experiments demonstrate that HPSD significantly improves T2V performance while also delivering notable TI2V gains, effectively strengthening the model's base generation ability.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- S-Lab, Nanyang Technological University(新加坡南洋理工大学S-Lab实验室)
- University of Science and Technology of China(中国科学技术大学)
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新研究院)
- Shanghai AI Laboratory(上海人工智能实验室)
- Adobe Research(奥多比研究院)
- The Chinese University of Hong Kong(香港中文大学)
- CPII under InnoHK(InnoHK下属CPII机构)
机构由 AI 辅助整理,请以论文原文为准。