AI 中文总结
本研究提出紧凑统一图像生成模型Swift-Image,采用6B单流DiT等技术,实现领先综合性能,压缩后3B模型无明显损失,还总结了相关实用经验。
AI 中文摘要
我们提出了Swift-Image,这是一款用于文本到图像生成、单图像编辑及多图像编辑的紧凑统一模型。我们的目标是探究在有限计算预算下,通过系统的训练工程,一个相对小型的视觉生成器能被推动到何种程度。Swift-Image采用高效的6B单流DiT架构,以及从宽泛语义覆盖逐步演进至高分辨率、更强视觉质量、统一生成-编辑监督的渐进式训练流程。在训练后阶段,我们采用并行专家强化学习,随后进行多教师在线策略蒸馏,以缓解异质目标间的干扰。我们进一步通过提示增强器(Prompt Enhancer)将用户请求转换为与生成器对齐的视觉规范,从而将高层推理与像素级渲染解耦。为实现高效部署,结构剪枝与少步蒸馏生成了3B参数及加速变体。Swift-Image仅用6B参数、243K GPU训练小时,就在评估的开源模型中实现了领先的综合性能;压缩后的3B模型几乎无性能损失,而少步蒸馏则以显著更少的采样步骤进一步提升了综合编辑性能。我们的研究还总结了架构、数据课程、训练后、提示增强及模型压缩方面的实用经验。
英文摘要
We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.
Comments28 pages, 11 figures