arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CrossDistill:通过轨迹级混合少步蒸馏平衡质量与多样性

CrossDistill: Balancing Quality and Diversity via Trajectory-Level Hybrid Few-Step Distillation

Yuxi Liu, Haoyu Li, Yixiang Cai, Tengxu Sun, Zekun Zhang, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kun Yuan, Kai Zhang

arXiv 2609.14725首次发表:更新:

发表机构

Peking University; Melon Group; Tsinghua University; Alibaba Group(北京大学; Melon集团; 清华大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CrossDistill提出轨迹级混合蒸馏框架,通过噪声区间划分与交叉耦合平衡质量与多样性,在文本到视频扩散模型上扩展了少步质量-多样性前沿。

AI 中文摘要

少步蒸馏加速了扩散模型,但必须在多样性与保真度之间取得平衡:基于轨迹的蒸馏保留了模式覆盖,而分布匹配使样本更锐利但可能降低多样性。我们表明,这种张力可以以依赖于噪声区间的方式加以利用:高噪声步骤主要决定全局模式,而低噪声步骤细化局部细节。我们提出CrossDistill,一种轨迹级混合蒸馏框架,它在交叉点处分割采样轨迹,在高噪声区间应用轨迹保持目标,在低噪声区间应用分布匹配目标,并通过交叉状态耦合这两个阶段。与损失级混合相比,并与训练时两阶段方案互补,CrossDistill沿噪声轴显式分配互补目标,从而在局部统计被锐化之前保留全局分支。CrossDistill是一种噪声级调度策略:PCM和DMD是可插拔的实例,而噪声划分、交叉耦合和目标排序是关键设计要素。在文本到视频扩散模型上的实验以及定性的图像到视频结果表明,CrossDistill扩展了少步质量-多样性前沿,在保持种子级变异的同时实现了有竞争力的视觉保真度。

英文摘要

Few-step distillation accelerates diffusion models but must balance diversity and fidelity: trajectory-based distillation preserves mode coverage, while distribution matching sharpens samples but can reduce diversity. We show that this tension can be exploited in a noise-regime-dependent way: high-noise steps largely determine global modes, whereas low-noise steps refine local details. We propose CrossDistill, a trajectory-level hybrid distillation framework that splits the sampling trajectory at a crossover point, applies a trajectory-preserving objective on the high-noise interval and a distribution-matching objective on the low-noise interval, and couples the two stages through the crossover state. In contrast to loss-level mixing, and complementarily to training-time two-stage recipes, CrossDistill explicitly assigns complementary objectives along the noise axis, so that global branching is preserved before local statistics are sharpened. CrossDistill is a noise-level scheduling policy: PCM and DMD are plug-in instantiations, while the noise partition, crossover coupling, and objective ordering are the key design elements. Experiments on text-to-video diffusion models and qualitative image-to-video results show that CrossDistill expands the few-step quality-diversity frontier, retaining seed-level variation while achieving competitive visual fidelity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑