AI 中文总结
介绍开源统一多模态模型Boogu-Image-0.1,通过改进模型理解、数据质量和训练管道等,结合智能推理扩展,提升生成和编辑性能,在标准基准测试中表现出色,成本低,还分享实践讨论并开源代码等推动相关生态发展。
AI 中文摘要
我们介绍了Boogu-Image-0.1,一个开源统一多模态理解与生成模型家族,包括Base、Turbo、Edit和Edit-Turbo变体。它在高质量文本到图像生成、快速推理、基于指令的编辑和双语(汉英)文本渲染方面表现出色。像Nano-Banana-Pro和GPT-Image-2等闭源多模态系统通过系统级集成而非单个模型取得强大性能,但其内部做法大多未公开。本文表明,在模型理解、数据质量和训练管道方面的针对性改进,加上推理时的智能扩展,即使在计算预算受限的情况下也能大幅提升生成和编辑性能。综合评估显示,Boogu-Image-0.1在标准基准测试中始终匹配或超越其他开源模型,结果接近领先的闭源系统。值得注意的是,这仅用了2.086亿张独特图像,基础模型的理论训练成本仅约40万美元。我们分享了对更广泛研究社区有价值的实践讨论,并在Apache 2.0下发布权重、代码和方法,以推动统一多模态理解与生成的开放生态系统发展。
英文摘要
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.