DART:面向少步视频扩散模型中免训练LoRA复用的蒸馏感知重参数化
DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models
- University of Electronic Science and Technology of China(电子科技大学)
- Tsinghua University(清华大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出免训练方法DART,结合低秩坐标传输与目标调度响应校准,提升少步视频扩散模型中LoRA复用质量,实验显示质量与功能保留显著改善。
AI中文摘要:
步骤蒸馏降低了视频生成的成本,但复用为更长轨迹训练的LoRA可能会改变其功能效果或降低目标质量。静态参数兼容性为这一问题提供了一个视角;我们的观察表明,在缩短的去噪调度下,相似的测量几何可以与不同的适配器行为共存。我们提出了DART,一种免训练方法,它结合了低秩坐标传输与基于前向评估的目标调度响应校准,且无需源训练视频。在四步Wan2.2目标上,DART-F将联合质量得分从0.9029提升至0.9227,并将宏观功能保留从-0.4644改变为+0.1349。组件分析表明,校准贡献了大部分质量提升,而坐标传输在与校准结合时提供了互补的增益。适配器级结果揭示了某些适配器的积极功能效果,以及另一些适配器的强衰减和减少的负面功能效果。对另外两个目标的评估显示了相同的总体趋势。这些结果促使通过功能保留和负迁移避免来联合评估蒸馏模型的LoRA复用,而不假设每个适配器都能恢复。
英文摘要:
Few-step distillation reduces the inference cost of image-to-video generation, but directly reusing LoRA adapters trained for long denoising trajectories can weaken their intended effects and degrade video quality. We observe that adapters with similar measured static parameter geometry can behave differently under the shortened target schedule, motivating response-aware transfer. We propose DART, a training-free reparameterization method that transports source LoRAs into aligned coordinates through a low-rank distillation bridge. Paired forward evaluations measure channel-level incremental responses under the target schedule to fit coefficients that calibrate response direction, source-relative magnitude, and timestep allocation. Coordinate transport establishes update directions, while calibration adapts their contributions, requiring neither source training videos nor backpropagation. Experiments across multiple distilled I2V models demonstrate improved generation quality and aggregate functional retention over direct reuse. Further analyses show that coordinate transport complements response calibration, with adapter-level benefits encompassing both functional preservation and reduced negative transfer.