发表机构
Microsoft Mage Team(微软Mage团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大规模视觉生成器成本高的问题,提出Mage-Flow,由Mage-VAE和原生分辨率多模态扩散Transformer组成。通过协同设计实现高效文本到图像生成及编辑,开发完整模型家族,Turbo变体在高分辨率下生成和编辑高效,性能有竞争力。
AI 中文摘要
大规模视觉生成器功能日益强大,但训练、微调及部署成本高昂。我们引入了Mage-Flow,这是一个紧凑的4B规模生成堆栈,用于高效的文本到图像生成和基于指令的图像编辑。该堆栈由两个协同设计的组件构建而成:Mage-VAE,一个轻量级高保真潜在分词器;以及一个通过整流流匹配训练的原生分辨率多模态扩散Transformer。Mage-VAE使用一步扩散式编码和解码以及锚定潜在正则化,在保持强大公共VAE重建质量的同时,将分词成本降低了一个多数量级。结合原生分辨率打包和堆栈级CUDA内核融合,该堆栈支持灵活分辨率训练,并将端到端训练吞吐量提高了约2.5倍。在此基础上,我们开发了一个完整的模型家族,包括用于生成和编辑的基础、RL对齐和Turbo变体。扩散-NFT改善了提示跟随、文本渲染、美学质量和编辑保真度,而通过对抗性感知指导的少步蒸馏产生了用于低延迟推理的4步Turbo模型。尽管规模紧凑,Mage-Flow和Mage-Flow-Edit在标准生成和编辑基准测试中实现了有竞争力的性能。更重要的是,Turbo变体使高分辨率生成和编辑在交互使用中变得切实可行:在单个NVIDIA A100 GPU上以1024^2分辨率,Mage-Flow-Turbo在0.59秒内生成图像,Mage-Flow-Edit-Turbo在1.02秒内编辑图像,同时保持较小的内存占用。这些结果表明,仔细的分词器-主干-系统协同设计可以在一个高效的4B模型家族中实现强大的高分辨率生成和编辑。
英文摘要
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about $2.5\times$. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at $1024^2$ resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.