从模型到系统:高效多模态学习综合综述
From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning
浏览论文内容
中文总结 AI 辅助
本综述提出首个从模型到系统的高效多模态学习分类法,涵盖模型、算法、系统三层,探讨跨层协同设计及效率-效用-隐私权衡,并展望自调节智能的未来方向。
中文摘要 AI 辅助
多模态模型的快速扩展暴露了计算、内存和部署方面的巨大瓶颈,催化了高效多模态学习(EML)作为关键研究前沿的兴起。尽管取得了密集进展,但对效率在学习栈中体现在何处、如何体现以及为何体现的连贯理解仍然支离破碎。本综述通过引入首个结构化的从模型到系统的分类法,系统化了EML领域。我们从300多篇开创性工作中提炼见解,将其分为三个层次——模型、算法和系统——分别解决架构精简、执行优化和硬件感知编排问题。超越纯粹的分类回顾,我们提供了这些层之间垂直协同的方法论综合,阐明了跨层协同设计如何促进基本的“效率-效用-隐私”权衡。通过多模态大语言模型(MLLMs)的整合案例研究,我们追溯了该领域从初始结构调整到现代全栈资源编排的演进轨迹。此外,我们为不同领域提供了整体性讨论和特定应用优化蓝图,并提出了向自我调节智能的范式转变,其中效率是模型基本设计的内在涌现属性,而非事后约束。最后,我们提出了将定义EML研究轨迹的开放挑战和未来方向。本综述为不仅高性能和可泛化,而且原生高效并准备普遍部署的多模态系统建立了结构化框架。持续更新版本可在https://this https URL获取。
英文摘要
The rapid expansion of multimodal models has surfaced formidable bottlenecks in computation, memory, and deployment, catalyzing the rise of Efficient Multimodal Learning (EML) as a pivotal research frontier. Despite intensive progress, a cohesive understanding of what, how, and where efficiency is manifested across the learning stack remains fragmented. This survey systematizes the EML landscape by introducing the first structured, model-to-system taxonomy. We distill insights from over 300 seminal works into three hierarchical levels--model, algorithm, and system--addressing architectural parsimony, execution refinement, and hardware-aware orchestration, respectively. Moving beyond a purely categorical review, we offer a methodological synthesis of the vertical synergies between these layers, elucidating how cross-layer co-design contributes to the fundamental "Efficiency-Utility-Privacy" trade-off. Through an integrative case study of Multimodal Large Language Models (MLLMs), we trace the field's evolutionary trajectory from initial structural adjustments to modern full-stack resource orchestration. Furthermore, we provide a holistic discussion and application-specific optimization blueprints for diverse domains and posit a paradigm shift toward self-regulating intelligence, where efficiency is an intrinsic, emergent property of the model's fundamental design rather than a post-hoc constraint. Finally, we present open challenges and future directions that will define the trajectory of EML research. This survey establishes a structured framework for multimodal systems that are not only high-performing and generalizable but natively efficient and ready for ubiquitous deployment. A continuously updated version is available at https://github.com/pwang322/Efficient-Multimodal-Learning-Survey.
发表机构
- University of Pittsburgh(匹兹堡大学)
- Stanford University(斯坦福大学)
- UCSD(加州大学圣地亚哥分校)
- University of Arizona(亚利桑那大学)
- Georgia Tech(佐治亚理工学院)
- Carnegie Mellon University(卡内基梅隆大学)
- UNC Chapel Hill(北卡罗来纳大学教堂山分校)
- UC Merced(加州大学默塞德分校)
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。