arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2405.16640cs.AIcs.CLcs.CVcs.MM

多模态大语言模型的数据中心视角综述

A Survey of Multimodal Large Language Model from A Data-centric Perspective

  • Hong Kong University of Science and Technology(香港科技大学)
  • Apple(苹果公司)
  • Peking University(北京大学)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianyi Bai, Hao Liang, Binwang Wan, Yanran Xu, Xi Li, Shiyu Li, Ling Yang, Bozhou Li, Yifan Wang, Bin Cui, Ping Huang, Jiulong Shan, Conghui He, Binhang Yuan, Wentao Zhang

更新

AI总结:

本文从数据中心视角系统综述多模态大语言模型,涵盖数据准备、评估方法及未来方向,旨在推动该领域研究。

AI中文摘要:

多模态大语言模型(MLLMs)通过整合和处理来自多种模态的数据(包括文本、视觉、音频、视频和3D环境)来增强标准大语言模型的能力。数据在模型的开发和优化中起着关键作用。在本综述中,我们从数据中心视角全面回顾了关于MLLMs的文献。具体来说,我们探讨了在MLLMs的预训练和适应阶段准备多模态数据的方法。此外,我们分析了数据集的评估方法,并回顾了用于评估MLLMs的基准。我们的综述还概述了潜在的未来研究方向。这项工作旨在为研究人员提供对MLLMs数据驱动方面的详细理解,促进该领域的进一步探索和创新。

英文摘要:

Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field.

↑