FlowMap-OPD:用于少步流图生成器在线策略蒸馏的展开-核分离
FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
FlowMap-OPD提出在线策略蒸馏框架,分离学生状态获取与分布比较,通过流-速度一致性实现少步流图生成器高效训练,在ImageNet和文本到图像任务中优于现有方法。
AI中文摘要:
少步流图生成器,包括MeanFlow和一致性模型,通过长距离传输实现高效采样,但其在线策略蒸馏仍未被充分探索。我们引入FlowMap-OPD,一种在线策略蒸馏框架,将学生状态获取与教师-学生分布比较分离。基于状态边缘的公式确立了这种分离,而流-速度一致性将局部监督与部署的长距离映射联系起来。在该框架内,我们开发了流图、诱导速度和瞬时速度分布监督,每种监督均与单独指定的原生流图展开配对。跨三种教师奖励的跨容量ImageNet实验表明,具有独立可调学生一致性的瞬时速度分布监督是最有效的选择。在文本到图像实验中,FlowMap-OPD展示了强大的多专家整合能力,并在任务性能和收敛速度上超越了多奖励Flow-Map GRPO。
英文摘要:
Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.