arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mage-Flow:用于图像生成和编辑的高效原生分辨率基础模型

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu

arXiv 2607.19064首次发表:更新:

发表机构

Microsoft Mage Team(微软Mage团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大规模视觉生成器成本高的问题,提出Mage-Flow,由Mage-VAE和原生分辨率多模态扩散Transformer组成。通过协同设计实现高效文本到图像生成及编辑,开发完整模型家族,Turbo变体在高分辨率下生成和编辑高效,性能有竞争力。

AI 中文摘要

大规模视觉生成器功能日益强大,但训练、微调及部署成本高昂。我们引入了Mage-Flow,这是一个紧凑的4B规模生成堆栈,用于高效的文本到图像生成和基于指令的图像编辑。该堆栈由两个协同设计的组件构建而成:Mage-VAE,一个轻量级高保真潜在分词器;以及一个通过整流流匹配训练的原生分辨率多模态扩散Transformer。Mage-VAE使用一步扩散式编码和解码以及锚定潜在正则化,在保持强大公共VAE重建质量的同时,将分词成本降低了一个多数量级。结合原生分辨率打包和堆栈级CUDA内核融合,该堆栈支持灵活分辨率训练,并将端到端训练吞吐量提高了约2.5倍。在此基础上,我们开发了一个完整的模型家族,包括用于生成和编辑的基础、RL对齐和Turbo变体。扩散-NFT改善了提示跟随、文本渲染、美学质量和编辑保真度,而通过对抗性感知指导的少步蒸馏产生了用于低延迟推理的4步Turbo模型。尽管规模紧凑,Mage-Flow和Mage-Flow-Edit在标准生成和编辑基准测试中实现了有竞争力的性能。更重要的是,Turbo变体使高分辨率生成和编辑在交互使用中变得切实可行:在单个NVIDIA A100 GPU上以1024^2分辨率,Mage-Flow-Turbo在0.59秒内生成图像,Mage-Flow-Edit-Turbo在1.02秒内编辑图像,同时保持较小的内存占用。这些结果表明,仔细的分词器-主干-系统协同设计可以在一个高效的4B模型家族中实现强大的高分辨率生成和编辑。

英文摘要

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about $2.5\times$. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at $1024^2$ resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑