arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态大语言模型的演进安全态势:新兴威胁与防护措施综述

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang

arXiv 2608.07535首次发表:更新:

发表机构

University of Alabama at Birmingham; NVIDIA; Penn State University; Columbia University; University of Missouri-Kansas City; Florida State University; Auburn University(阿拉巴马大学伯明翰分校; 英伟达公司; 宾夕法尼亚州立大学; 哥伦比亚大学; 密苏里大学堪萨斯分校; 佛罗里达州立大学; 奥本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述针对多模态大语言模型(MLLMs)的新型安全威胁,提出多模态安全威胁分类法,梳理安全策略进展,探讨未来安全机制的研究方向。

AI 中文摘要

多模态大语言模型(MLLMs)通过模态对齐与融合整合异构模态,实现更强的理解与推理能力。但这种架构转变重塑了机器学习的安全态势,模型复杂度提升与跨模态交互催生了新型威胁,包括模态集成受损、模态对齐偏差及融合安全风险,反映出威胁建模从单模态假设的转变。这些转变对安全方案施加了现有单模态学习框架未覆盖的新约束。受这些挑战驱动,本综述系统分析MLLMs的演进安全态势:首先提出基于多模态的安全威胁分类法,分析威胁模型的转变,涵盖对抗攻击、数据投毒、越狱攻击及幻觉;随后总结更新后的安全假设,并据此梳理MLLM安全策略的最新进展;最后探讨开放挑战与未来方向,为开发更具原则性与可扩展性的多模态系统安全机制提供参考。

英文摘要

Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused safety risks, reflecting shifts in threat modeling beyond uni-modal assumptions. These shifts, in turn, impose new constraints on safety solutions not captured by existing frameworks rooted in uni-modal learning. Motivated by these challenges, this survey provides a systematic analysis of the evolving safety landscape of MLLMs. We first propose a multimodal grounded taxonomy of safety threats and analyze shifts in threat models, covering adversarial attacks, data poisoning, jailbreaks, and hallucinations. We then summarize updated safety assumptions and organize recent advances in MLLM safety strategies accordingly. Finally, we discuss open challenges and future directions to inform the development of more principled and scalable safety mechanisms for multimodal systems.

CommentsAccepted at the ICLR 2026 Workshop on Principled Design for Trustworthy AI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑