LLaDA-Image:基于完全开放训练方案构建强大图像生成器
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
浏览论文内容
中文总结 AI 辅助
该研究提出LLaDA-Image统一框架,结合6B DiT与冻结视觉-语言模块,经仅图像预训练等优化,在Qwen-Image-Bench创开源模型最优,还蒸馏出快速推理版本,发布相关资源。
中文摘要 AI 辅助
我们提出LLaDA-Image,这是一个统一框架,它将从头训练的60亿参数扩散Transformer(DiT)与基于LLaDA2.0-Mini扩散语言模型主干构建的冻结视觉-语言理解模块配对。与从一开始就严重依赖配对的图像-文本数据不同,我们首先通过仅图像的预训练和中间训练构建强大的视觉生成先验。生成流程包含2.2亿个样本,其中98个是真实图像。为实现高效且可扩展的优化,我们在整个DiT中使用无参数RMSNorm,同时结合Muon优化器。最终得到的统一模型能生成高度逼真的图像,同时准确遵循细粒度编辑指令。我们进一步将LLaDA-Image蒸馏为LLaDA-Image-Turbo,使其能在2至4个采样步骤内快速推理。在Qwen-Image-Bench上,LLaDA-Image在英文赛道和中文赛道分别取得53.53和53.38的总分,在两条赛道的开源模型中达到新的最优水平。为支持对强大且高效的生成模型的进一步研究,我们发布了模型权重、训练代码和详细方案。
英文摘要
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.
发表机构
- AGI Research Center, Inclusion AI(AGI研究中心,Inclusion AI)
机构由 AI 辅助整理,请以论文原文为准。