多模态智能体框架的基础与前沿:技术及应用综述
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
浏览论文内容
中文总结 AI 辅助
本综述针对多模态在智能体框架核心模块的作用展开系统分析,梳理了其架构整合方式、应用领域及效率权衡,为通用智能系统研究指明方向。
中文摘要 AI 辅助
大型语言模型(LLMs)的进展推动了智能体研究热潮,智能体具备推理、规划与行动的能力。该研究催生了以强大LLM为骨干,协调感知、记忆与决策的智能体框架。随着大型多模态模型(LMMs)的出现,这些系统可处理并整合图像、音频、视频等多种模态,提升了现实应用适用性。然而,尽管已有针对LLM智能体的综述,但近年来多模态在塑造智能体方面的作用尚未得到系统考察。本综述填补这一空白,通过分析多模态对智能体框架核心功能模块(感知、推理、规划、记忆、行动)的影响,从文本中心智能体追溯到多模态框架的演变,考察模态如何通过委托式、后期融合、前期融合架构整合,评估由具身感知与多模态推理催生的智能体行为。我们通过以模态为中心的分类法组织现有研究,将架构设计选择与智能体能力关联;还综述了机器人、图形用户界面(GUI)与网页导航、多媒体内容生成与编辑、长视频理解与检索等多应用领域的多模态智能体系统。除能力外,我们分析了这些场景下的性能,讨论了效率与可扩展性权衡,包括训练与推理成本、延迟及部署约束。通过聚焦多模态在智能体设计中的影响,我们旨在识别关键缺口,绘制通往稳健通用智能系统的路线图。
英文摘要
Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbones. With the advent of large multimodal models (LMMs), these systems can process and integrate diverse modalities, including images, audio, and video, thereby improving their real-world applicability. Yet, while surveys of LLM-based agents exist, the role of multimodality in shaping agency has not been systematically examined in recent years. This survey fills the gap by analyzing the impact of multimodality across the core functional modules of the agentic framework: perception, reasoning, planning, memory, and action. Using this lens, we trace the evolution from text-centric agents to multimodal frameworks, examine how modalities are integrated through delegated, late-fusion, and early-fusion architectures, and assess the emergence of agentic behaviors enabled by grounded perception and multimodal reasoning. We organize existing work through a modality-centric taxonomy that links architectural design choices to agent capabilities. Moreover, we review multimodal agentic systems across various application domains, including Robotics, GUI & Web Navigation, Multimedia Content Generation & Editing, and Long-form Video Understanding & Retrieval. Beyond capabilities, we analyze performance across these settings and discuss efficiency-scalability trade-offs, including training and inference costs, latency, and deployment constraints. By focusing on the impact of multimodality in agentic design, we aim to identify key gaps and chart a roadmap toward robust and general-purpose intelligent systems.