AI 中文总结
Any-OPD是首个针对任意潜在流匹配生成器对的策略上蒸馏框架,通过表示空间桥接实现异构模型间的蒸馏,以五分之一规模达到教师模型性能,解决了直接蒸馏无法训练的问题。
AI 中文摘要
策略上蒸馏中,教师模型会校正学生模型自身生成的样本,其前提是两个模型遵循相同的规则:拥有相同的VAE隐变量、匹配的架构以及共同的时间步网格。我们探究当这些条件全部不成立时会发生什么,例如当可用的最强教师模型与想要部署的学生模型来自不同模型家族时,发现标准方法无法解决该问题:教师隐变量无法作为异坐标系统中的目标;针对随机重绘局部细节的教师模型的逐像素损失会退化为模糊或发散;时间步索引在不匹配的调度中失去意义。我们提出Any-OPD,据我们所知,这是首个针对任意潜在流匹配生成器对的策略上蒸馏框架。Any-OPD将教师模型纯粹视为黑盒采样器,并在一个关键点连接两个模型:一个冻结的、与模型无关的视觉表示,在该表示中比较它们独立解码的输出,从而规避对隐变量、特征或架构的所有假设。通过匹配连续噪声水平而非步索引来恢复轨迹对应关系,以及一个简短的锚定阶段——在此阶段,教师样本通过学生自身的VAE重新编码——确保策略上梯度衡量样本质量而非域不匹配。将12B规模的FLUX.1-dev蒸馏为2.5B规模的SD3.5-Medium时,Any-OPD将学生模型的PickScore从0.846提升至0.884,HPSv3从9.12提升至10.97,达到教师模型的性能而规模仅为其五分之一,而直接隐变量回归则完全无法训练。
英文摘要
On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid. We ask what happens when none of this holds, as when the strongest teacher available and the student one wishes to deploy come from different model families, and find that the standard recipes have no answer: teacher latents cannot serve as targets in a foreign coordinate system, per-pixel losses against a teacher that stochastically re-draws local detail degenerate into blur or divergence, and timestep indices lose their meaning across mismatched schedules. We present Any-OPD, to our knowledge the first framework for on-policy distillation between arbitrary pairs of latent flow-matching generators. Any-OPD treats the teacher purely as a black-box sampler and connects the two models at exactly one point: a frozen, model-agnostic vision representation in which their independently decoded outputs are compared, sidestepping every assumption about latents, features, or architecture. Trajectory correspondence is recovered by matching continuous noise levels instead of step indices, and a brief anchoring phase, in which teacher samples are re-encoded through the student's own VAE, ensures the on-policy gradient measures sample quality rather than domain mismatch. Distilling the 12B FLUX.1-dev into the 2.5B SD3.5-Medium, Any-OPD lifts the student's PickScore from 0.846 to 0.884 and HPSv3 from 9.12 to 10.97, rivaling the teacher at a fifth of its size, where direct latent regression fails to train at all.