多模态大语言模型评估综述
A Survey on Evaluation of Multimodal Large Language Models
- College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文系统综述了多模态大语言模型(MLLMs)的评估方法,从背景、评估内容(感知、推理及特定领域应用等)、评估基准和评估步骤与指标四个方面进行了全面梳理,强调评估对推进该领域发展的关键作用。
AI中文摘要:
多模态大语言模型(MLLMs)通过将强大的大语言模型(LLMs)与各种模态编码器(例如视觉、音频)相集成,模拟人类的感知和推理系统,将LLMs定位为“大脑”,将各种模态编码器定位为感觉器官。这种框架赋予MLLMs类人的能力,并暗示了迈向通用人工智能(AGI)的潜在途径。随着GPT-4V和Gemini等全能MLLMs的出现,大量评估方法被开发出来,以在不同维度上评估它们的能力。本文对MLLM评估方法进行了系统全面的回顾,涵盖以下关键方面:(1)MLLMs的背景及其评估;(2)“评估什么”,基于所评估的能力对现有MLLM评估任务进行回顾和分类,包括通用多模态识别、感知、推理和可信度,以及特定领域应用,如社会经济、自然科学和工程、医疗用途、AI智能体、遥感、视频和音频处理、3D点云分析等;(3)“在哪里评估”,将MLLM评估基准总结为通用和特定基准;(4)“如何评估”,回顾并说明MLLM评估步骤和指标;我们的总体目标是为MLLM评估领域的研究人员提供有价值的见解,从而促进更强大、更可靠的MLLMs的发展。我们强调,评估应被视为一门关键学科,对推进MLLMs领域至关重要。
英文摘要:
Multimodal Large Language Models (MLLMs) mimic human perception and reasoning system by integrating powerful Large Language Models (LLMs) with various modality encoders (e.g., vision, audio), positioning LLMs as the "brain" and various modality encoders as sensory organs. This framework endows MLLMs with human-like capabilities, and suggests a potential pathway towards achieving artificial general intelligence (AGI). With the emergence of all-round MLLMs like GPT-4V and Gemini, a multitude of evaluation methods have been developed to assess their capabilities across different dimensions. This paper presents a systematic and comprehensive review of MLLM evaluation methods, covering the following key aspects: (1) the background of MLLMs and their evaluation; (2) "what to evaluate" that reviews and categorizes existing MLLM evaluation tasks based on the capabilities assessed, including general multimodal recognition, perception, reasoning and trustworthiness, and domain-specific applications such as socioeconomic, natural sciences and engineering, medical usage, AI agent, remote sensing, video and audio processing, 3D point cloud analysis, and others; (3) "where to evaluate" that summarizes MLLM evaluation benchmarks into general and specific benchmarks; (4) "how to evaluate" that reviews and illustrates MLLM evaluation steps and metrics; Our overarching goal is to provide valuable insights for researchers in the field of MLLM evaluation, thereby facilitating the development of more capable and reliable MLLMs. We emphasize that evaluation should be regarded as a critical discipline, essential for advancing the field of MLLMs.