arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04349cs.CV

Poly-OPD:用于能力可选流模型的异构多教师在线策略蒸馏

Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models

Siming Fu, Haojun Xu, Ruizhe He, Zheming Fu, Hualiang Wang, Jie Huang, Xiaoxiao Ma, Mingchen Zhong, Weihu Huang, Xiaoxuan He, Linjiang Huang, Si Liu

AI总结:

Poly-OPD框架通过异构多教师在线策略蒸馏,将FLUX.1-dev和Z-Image的互补优势整合到25亿参数的SD3.5-Medium模型,提升了文本到图像生成的GenEval与DrawBench指标。

AI中文摘要:

主流开源文本到图像模型通常具有互补优势:一个可能在偏好对齐的美学表现上领先,而另一个则更忠实地遵循组合指令。然而,它们的自动编码器和噪声调度器存在差异,导致跨模型转移这些优势十分困难。本文提出Poly-OPD,这一框架可将异构教师的互补优势整合到单个紧凑的流匹配学生模型中。为桥接不同教师不兼容的潜在空间,Poly-OPD通过像素桥执行在线策略蒸馏:每个学生生成的图像会被选定教师的编码器重新编码,并在教师噪声调度器下按幅度匹配的噪声水平进行优化;生成的目标会在冻结的DINOv2空间中进一步与学生匹配,实现跨不兼容潜在空间的监督。为保留互补能力且避免跨教师干扰,Poly-OPD采用梯度兼容性诊断来组织其适配器:注意力LoRA模块在教师间共享,而前馈适配器则为教师专属。蒸馏过程中,差距感知课程会将更多训练投入学生仍落后于教师的组合类别;随着差距缩小,训练会转向剩余差距更大的类别。通过将FLUX.1-dev和Z-Image蒸馏至25亿参数的SD3.5-Medium学生模型,Poly-OPD将GenEval指标从67.3提升至73.3,超越两个更大的教师模型,并将DrawBench HPSv3从9.34提升至11.35,在可切换模型中整合了两种优势。

英文摘要:

Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to transfer these strengths across models. In this paper, we present Poly-OPD, a framework that can consolidate complementary strengths of heterogeneous teachers into a single compact flow-matching student. To bridge the incompatible latent spaces of different teachers, Poly-OPD performs on-policy distillation through a pixel bridge. Each student-generated image is re-encoded by a selected teacher's encoder and refined from a noise level matched by magnitude under the teacher's noise schedule. The resulting target is further matched to the student in frozen DINOv2 space, enabling supervision across incompatible latent spaces. To retain complementary capabilities without cross-teacher interference, Poly-OPD uses a gradient compatibility diagnostic to organize its adapters: attention LoRA modules are shared across teachers, whereas feed-forward adapters remain teacher-specific. During distillation, a gap-aware curriculum devotes more training to compositional categories where the student still falls short of the teacher. As each gap narrows, training shifts toward categories with larger remaining gaps. By distilling FLUX.1-dev and Z-Image into a 2.5B SD3.5-Medium student, Poly-OPD improves GenEval from 67.3 to 73.3, surpassing both larger teachers, and raises DrawBench HPSv3 from 9.34 to 11.35, consolidating both strengths within a switchable model.

↑