发表机构
Machine Learning and Data Science Unit; Okinawa Institute of Science and Technology Graduate University(机器学习与数据科学部; 冲绳科学技术大学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对流映射蒸馏,提出几何感知时间重参数化方法,通过分配更多学生时间给高法向加速度区域,在保持教师路径的同时提升一步和少步生成质量,实验验证其有效性。
AI 中文摘要
流映射蒸馏通过学习预训练生成式常微分方程的有限时间转移,实现一步和少步生成。我们研究改变教师模型的时间参数化是否能使这些转移更容易学习。基于大法向加速度的轨迹段更难蒸馏的假设,我们提出一种几何感知的时间重参数化方法,该方法将更多学生时间分配给这些区域,同时保持教师模型的几何路径和终端分布。我们在适当假设下推导出一个共享时钟,该时钟均衡总体法向加速度统计量,并利用跨教师轨迹的鲁棒正则化估计构建实用近似。我们将该时钟纳入拉格朗日流映射蒸馏中,使用变换后的时间坐标来条件化学生模型。该时钟在蒸馏前估计一次,既不需要重新训练教师模型,也不需要额外的学生参数或推理时网络评估。在合成数据、CIFAR-10和CelebA-64上的实验表明,在匹配推理预算下,与恒等时间蒸馏相比,样本质量有所提高,包括一步和两步图像生成的改进。一步生成中的增益(其中无法调整中间采样时间)凸显了蒸馏过程中时间重参数化的优势。
英文摘要
Flow-map distillation enables one- and few-step generation by learning finite-time transitions of a pretrained generative ODE. We investigate whether changing the teacher's time parameterization can make these transitions easier to learn. Motivated by the hypothesis that trajectory segments with large normal acceleration are harder to distill, we propose a geometry-aware time reparameterization that allocates more student time to these regions while preserving the teacher's geometric paths and terminal distribution. We derive a shared clock that equalizes a population normal-acceleration statistic under suitable assumptions, and construct a practical approximation from robust, regularized estimates across teacher trajectories. We incorporate this clock into Lagrangian flow-map distillation, using the transformed time coordinate to condition the student. The clock is estimated once before distillation and requires neither teacher retraining nor additional student parameters or inference-time network evaluations. Experiments on synthetic data, CIFAR-10, and CelebA-64 show improved sample quality over identity-time distillation at matched inference budgets, including improvements in one- and two-step image generation. The gains in one-step generation, where no intermediate sampling times can be adjusted, highlight the benefits of time reparameterization during distillation.