arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从语料到协同进化的能力:面向通用图像生成的以能力为中心的数据设计

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Zhengrui Chen, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Yuan Wang, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen

arXiv 2608.18076首次发表:更新:

发表机构

Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出能力驱动的数据基础设施,构建三类监督引擎,整理大规模图像语料,训练多模态扩散模型,在CPI-Bench等评估中展现良好生成与迁移能力。

AI 中文摘要

大规模图像生成受益于数据规模、质量、重平衡及重标注的进展,但传统流程通常孤立地优化特定任务的数据集。核心挑战不仅在于如何整理每个特定任务的语料库,还在于如何根据生成能力间的依赖关系组织异构监督。我们提出一种**能力驱动的数据基础设施**,它将特定能力的监督构建与符合能力对齐的课程调度相结合。其三个专用且可互操作的数据引擎为文本-图像 grounding、图像间转换及图像-知识关联构建互补的关系型监督,同时标注专家在各任务及粒度级别对齐文本到图像生成(T2I)与编辑监督。多阶段课程沿能力获取的依赖顺序协同演进任务组成、视觉概念分布、数据质量及图像分辨率,而能力感知评估通过针对性检索、专家构建及感知差距的重采样形成闭环。在大规模应用中,该框架整理出4.4亿张图像的T2I语料库、1.2亿个编辑对及超2700万个图像-实体对。基于此基础设施,我们从头开始训练了两个规模分别为30亿和60亿参数的多模态扩散模型。我们在CPI-Bench上开展定量评估,并在多种文本到图像生成及编辑场景中开展定性评估。实验结果显示该模型具备广泛的视觉覆盖范围、通用的渲染能力及跨生成能力的有效迁移性。

英文摘要

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

Comments19 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑