arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02290cs.CV

DisciplineGen-1M:用于多学科视觉生成与编辑的大规模数据集

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

Zhaokai Wang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Yiguo He, Mohan Zhang, Leyao Gu, Yan Li, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Xuanhe Zhou, Zhihang Zhong, Xue Yang

首次发表
浏览论文内容

中文总结 AI 辅助

提出百万级多学科数据集DisciplineGen-1M,涵盖10个学科,通过可扩展框架构建,并基于此开发学科感知推理生成模型,在学科和通用推理基准上显著提升。

中文摘要 AI 辅助

最近的图像生成和编辑模型可以产生视觉上吸引人的自然图像,但当目标图像是知识密集型图表时,其正确性依赖于学科概念、符号结构和精确的空间关系,这些模型仍然不可靠。我们引入了DisciplineGen-1M,一个百万级的多学科数据集,支持文本到图像生成和图像编辑。它包含120万个样本,涵盖数学、物理、化学、生物学、地理、计算机科学、经济学、历史、音乐和体育。为了构建该数据集,我们设计了一个可扩展的框架,结合了矢量图形渲染、基于OCR的编辑、策划的程序化合成和大规模文本到图像过滤。这些流程生成标题、编辑指令、结构化注释以及具有可控语义差异的配对图像。基于DisciplineGen-1M,我们进一步引入了一个学科感知的推理生成模型,用于文本到图像生成和图像编辑。在学科相关基准GenExam和GRADE上的实验表明,与开源基线相比有显著改进,而在通用推理基准WISE和RISE上的评估进一步表明了更广泛的迁移。结果表明,大规模结构化学术视觉数据是将图像生成从美学合理性转向可验证的知识基础视觉创作的关键因素。我们将公开发布我们的数据集、模型和数据整理流程的源代码,以确保可重复性并惠及未来研究。

英文摘要

Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose correctness depends on disciplinary concepts, symbolic structure, and precise spatial relations. We introduce DisciplineGen-1M, a million-scale multidisciplinary dataset that supports text-to-image generation and image editing. It contains 1.2M samples spanning mathematics, physics, chemistry, biology, geography, computer science, economics, history, music, and sports. To construct the dataset, we design a scalable framework that combines structured rendering, OCR-based editing, specialized programmatic synthesis, and large-scale text-to-image filtering. These pipelines produce captions, editing instructions, structured annotations, and paired images with controllable semantic differences. Building on DisciplineGen-1M, we further introduce a discipline-informed reasoning-generation model for both text-to-image generation and image editing. Experiments on discipline-related benchmarks, GenExam and GRADE, show substantial improvements over open-source baselines, while evaluations on general reasoning-informed benchmarks, WISE and RISE, further indicate broader transfer. The results suggest that large-scale structured academic visual data is a key ingredient for moving image generation from aesthetic plausibility toward verifiable knowledge-grounded visual creation. We will publicly release our dataset, model, and source code of the data curation pipeline to ensure reproducibility and benefit future research.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • South China University of Technology(华南理工大学)
  • Xiamen University(厦门大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑