发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出能力驱动的数据基础设施,构建三类监督引擎,整理大规模图像语料,训练多模态扩散模型,在CPI-Bench等评估中展现良好生成与迁移能力。
AI 中文摘要
大规模图像生成受益于数据规模、质量、重平衡及重标注的进展,但传统流程通常孤立地优化特定任务的数据集。核心挑战不仅在于如何整理每个特定任务的语料库,还在于如何根据生成能力间的依赖关系组织异构监督。我们提出一种**能力驱动的数据基础设施**,它将特定能力的监督构建与符合能力对齐的课程调度相结合。其三个专用且可互操作的数据引擎为文本-图像 grounding、图像间转换及图像-知识关联构建互补的关系型监督,同时标注专家在各任务及粒度级别对齐文本到图像生成(T2I)与编辑监督。多阶段课程沿能力获取的依赖顺序协同演进任务组成、视觉概念分布、数据质量及图像分辨率,而能力感知评估通过针对性检索、专家构建及感知差距的重采样形成闭环。在大规模应用中,该框架整理出4.4亿张图像的T2I语料库、1.2亿个编辑对及超2700万个图像-实体对。基于此基础设施,我们从头开始训练了两个规模分别为30亿和60亿参数的多模态扩散模型。我们在CPI-Bench上开展定量评估,并在多种文本到图像生成及编辑场景中开展定性评估。实验结果显示该模型具备广泛的视觉覆盖范围、通用的渲染能力及跨生成能力的有效迁移性。
英文摘要
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.
Comments19 pages, 10 figures