arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Swift-Image:探索紧凑统一图像生成模型的性能前沿

Exploring the Performance Frontier of Compact Unified Image Generation Models

Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu, Hao Yan, Zhengze Xu, Yuhang Yu, Yongchao Du, Xingjian Wang, Jun Zheng, Qinye Zhou, Yaqi Cai, Zhengrui Chen, Chao Lin, Yefeng Shen, Yuan Wang, Zhengtao Wu, Ge Wu, Xiaoli Xu, Denghui Yang, Huayu Zhang, Mingzhou Zhang, Mengting Chen

arXiv 2608.20334首次发表:更新:

AI 中文总结

本研究提出紧凑统一图像生成模型Swift-Image,采用6B单流DiT等技术,实现领先综合性能,压缩后3B模型无明显损失,还总结了相关实用经验。

AI 中文摘要

我们提出了Swift-Image,这是一款用于文本到图像生成、单图像编辑及多图像编辑的紧凑统一模型。我们的目标是探究在有限计算预算下,通过系统的训练工程,一个相对小型的视觉生成器能被推动到何种程度。Swift-Image采用高效的6B单流DiT架构,以及从宽泛语义覆盖逐步演进至高分辨率、更强视觉质量、统一生成-编辑监督的渐进式训练流程。在训练后阶段,我们采用并行专家强化学习,随后进行多教师在线策略蒸馏,以缓解异质目标间的干扰。我们进一步通过提示增强器(Prompt Enhancer)将用户请求转换为与生成器对齐的视觉规范,从而将高层推理与像素级渲染解耦。为实现高效部署,结构剪枝与少步蒸馏生成了3B参数及加速变体。Swift-Image仅用6B参数、243K GPU训练小时,就在评估的开源模型中实现了领先的综合性能;压缩后的3B模型几乎无性能损失,而少步蒸馏则以显著更少的采样步骤进一步提升了综合编辑性能。我们的研究还总结了架构、数据课程、训练后、提示增强及模型压缩方面的实用经验。

英文摘要

We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.

Comments28 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑