arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PE-OPSD:通过在线策略自蒸馏将提示增强内化到流匹配模型中

PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation

Mingfeng Lin, Chengfei Cai, Lin Xu, Chengqian Ma, Yuxiang Wei, Liang Han

arXiv 2609.36638首次发表:更新:

发表机构

Harbin Institute of Technology (Shenzhen); Zhejiang University; Peking University(哈尔滨工业大学(深圳); 浙江大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PE-OPSD,通过在线策略自蒸馏将提示增强内化到流匹配模型中,在推理时无需额外延迟即可提升提示保真度,并保持推理效率。

AI 中文摘要

文生图用户经常提供简洁且不明确的提示,而生成模型则受益于详细的文本条件以实现可靠的指令遵循。现有系统通过提示增强器(PE)在推理时重写原始提示来弥合这一差距,这引入了额外的延迟,并使提示细化过程外置于生成器。相反,我们将增强提示视为特权训练信息,并探究其益处能否被内化。我们提出用于文生图流匹配模型的提示增强在线策略自蒸馏(PE-OPSD)。在训练期间,原始提示学生模型遵循其自身的生成轨迹,而增强提示教师模型在学生访问的状态处提供向量场目标。这种在线策略监督将增强提示所诱导的行为蒸馏到原始提示学生模型中,而无需额外的文本-图像对。在推理时,PE和教师模型均被移除,学生模型直接从原始提示生成。在多个模型家族、PE和基准测试中,PE-OPSD在评估的基线中实现了最强的整体提示保真度,产生了正向的整体视觉吸引力增益,并保持了基础模型的推理效率。

英文摘要

Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑