arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MLLM-DataEngine:闭合多模态指令微调数据生成的循环

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation

Zhiyuan Zhao, Bin Wang, Linke Ouyang, Yiqi Lin, Pan Zhang, Xiaoyi Dong, Jiaqi Wang, Conghui He

arXiv 2607.15299首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MLLM-DataEngine闭环系统,通过自适应坏例采样模块分析模型弱点,为GPT-4提供信息以生成高质量增量数据集,能有针对性且自动地提升MLLMs能力,有望成为MLLMs数据管理通用方案。

AI 中文摘要

在本文中,我们提出了MLLM-DataEngine,这是一种新颖的闭环系统,它连接了数据生成、模型训练和评估。在每个循环迭代中,MLLM-DataEngine首先根据评估结果分析模型的弱点,然后为下一次训练迭代生成合适的增量数据集,并迭代地增强模型能力。与以前与基准测试分开的指令微调数据集收集方法相比,MLLM-DataEngine具有更好的针对性,能更有效地提高多模态语言模型(MLLMs)的能力。首先,我们提出了一个自适应坏例采样模块,它可以根据基准测试结果有效地分析模型弱点,并灵活调整增量数据集的生成。其次,为了确保特定能力类型的高质量数据,向GPT-4提供最具代表性的上下文示例和丰富信息,这有助于GPT-4充分理解模型的弱点,并进一步保证生成的高质量数据。通过广泛的实验,我们发现MLLM-DataEngine可以在无需人工参与的情况下有针对性地、自动地提高MLLMs的能力。我们希望MLLM-DataEngine能成为后续MLLMs数据管理的通用解决方案。代码、数据和模型可在该https网址获取。

英文摘要

In this paper, we propose MLLM-DataEngine, a novel closed-loop system that bridges data generation, model training, and evaluation. Within each loop iteration, the MLLM-DataEngine first analyzes the weakness of the model based on the evaluation results, then generates a proper incremental dataset for the next training iteration, and enhances the model capability iteratively. Compared with previous instruction fine-tuning dataset collection methods which are separate from the benchmarking, MLLM-DataEngine shows better targeting and can improve MLLMs's capabilities more effectively. Firstly, we propose an Adaptive Bad-case Sampling module, which can effectively analyze model weakness based on the benchmarking results and adjust the generation of incremental datasets flexibly. Secondly, in order to ensure high-quality data for specific capability types, the most representative in-context examples and abundant information are provided to GPT-4, which helps GPT-4 fully comprehend the model's weakness and further guarantees high-quality generated data. Through extensive experiments, we find MLLM-DataEngine could boost the MLLMs capability in a targeted and automatic manner without human participants. We hope MLLM-DataEngine could be a general solution for the following MLLMs data curation. Code, data, and model are available at https://github.com/opendatalab/MLLM-DataEngine.

Comments6 pages, 4 figures, 7 tables; accepted by ICME 2026

Journal ref2025 IEEE International Conference on Multimedia and Expo (ICME), Nantes, France, 30 June 2025 - 04 July 2025

DOI:10.1109/ICME59968.2025.11208956

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑