从结构到协同:多模态大语言模型中视觉-语言感知范式演进综述
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
- School of Computer Science, Sichuan University(四川大学计算机学院)
- School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(北京大学深圳研究生院电子与计算机工程学院)
- Institute of Artificial Intelligence (TeleAI), China Telecom and Northwestern Polytechnical University(中国电信与西北工业大学人工智能研究院(TeleAI))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文系统综述多模态大语言模型中统一视觉-语言感知的范式演进,提出五阶段分类法,梳理各阶段代表性方法,并指出开放挑战与未来方向。
AI中文摘要:
多模态大语言模型(MLLMs)近期在统一视觉-语言理解与推理方面取得了显著进展,尤其是在OpenAI的O系列和DeepSeek的R系列模型引入后,推动了向感知中心智能的范式转变。然而,目前仍缺乏从真正统一的视觉-语言视角——即将视觉和语言视为不可分割的模态——来审视感知的系统性综述。现有综述往往碎片化,分别聚焦于视觉或语言,因此很少捕捉感知作为集成能力的跨模态演进。为填补这一空白,我们提出了首个关于MLLMs中统一视觉-语言感知的系统性综述。具体而言,我们(1)将MLLM感知形式化为一种类似于人类先天感知的内在、统一的视觉-语言能力,(2)引入一个五阶段分类法,追踪MLLM感知的范式演进,并调研每个阶段的代表性方法和里程碑,(3)识别开放挑战并勾勒出通向真正通用、统一多模态智能的有前景研究方向。我们希望我们的研究能为通向人工通用智能(AGI)的进一步创新提供基础理解和可操作路线图。
英文摘要:
Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-series, which have driven a paradigm shift toward perception-centric intelligence. However, there remains a lack of systematic surveys that examine perception from a truly unified vision-language perspective -- one that treats vision and language as an inseparable modality. Existing reviews are often fragmented, focusing separately on either vision or language, and thus rarely capture the cross-modal evolution of perception as an integrated capability. To bridge this gap, we present the first systematic survey of unified vision-language perception in MLLMs. Specifically, we (1) formalize MLLM perception as an intrinsic, unified vision-language capability analogous to human innate perception, (2) introduce a five-stage taxonomy tracing the paradigm evolution of MLLM perception and survey representative methods and milestones at each phase, and (3) identify open challenges and outline promising research directions toward truly general, unified multimodal intelligence. We hope our study will provide both a foundational understanding and an actionable roadmap to foster further innovation on the path toward artificial general intelligence (AGI).