发表机构
The Hong Kong University of Science and Technology; ATH, Alibaba Group; The Hong Kong University of Science and Technology (Guangzhou); Northwestern Polytechnical University; Zhejiang University(香港科技大学; 阿里巴巴集团 ATH; 香港科技大学(广州); 西北工业大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态指令数据集构建的人力成本高、现有方法模态支持有限及多轮指令生成困难的问题,提出UniData通用多模态指令生成流水线,构建含2万条9种模态的UniDataset,实现SOTA数据质量并提升其他多模态模型能力。
AI 中文摘要
多模态大语言模型(Multimodal Large Language Models, MLLMs)正被越来越多地应用于更广泛的现实场景。然而,由于人力成本高昂,为MLLMs创建高质量的多模态指令数据集仍是一项重大挑战。尽管一些方法提出生成指令数据,但它们往往在模态支持方面存在局限,且难以生成多轮指令。为解决这些问题,我们引入UniData,一种通用指令生成流水线,用于将简单的用户需求转化为多轮、多模态指令。具体而言,UniData首先将用户需求扩展为多个不同的事件;利用这些事件,UniData随后集成任意到任意的大模型以生成多模态指令;最后,UniData通过修正不相关和冗余的推理流程,利用多轮指令之间的相关性提升数据质量。为训练该流水线,我们还构建了UniDataset,一个包含9种模态共20000条条目的数据集,用于改进多模态生成。我们的实验表明,UniData在数据质量方面达到了SOTA性能,还可提升其他多模态模型的理解与生成能力。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they often face limitations in modality support and struggle with generating multi-round instructions. To address these problems, we introduce UniData, a universal instruction generation pipeline, to transform simple user requirements into multi-round, multimodal instructions. Specifically, UniData first expands user requirements into multiple diverse events. Using these events, UniData then integrates an any-to-any large model for multimodal instruction generation. Finally, UniData enhances data quality by correcting irrelevant and redundant inference flow, leveraging correlations between instruction rounds. To train this pipeline, we also build UniDataset, a dataset comprising 20,000 entries across nine modalities for improved multimodal generation. Our experiments demonstrate that UniData achieves SOTA performance in data quality and can also enhance the understanding and generation capabilities of other multimodal models.
CommentsAccepted by EMNLP 2026 Findings