arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MME-Safety:多模态大语言模型安全性的细粒度基准

MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

Yilian Shi, Yueming Lyu, Haoxiang Tan, Linzhuang Zou, Qihao Wang, Guihua Yu, Chenyang Si, Caifeng Shan

arXiv 2609.20850首次发表:更新:

发表机构

Nanjing University; Meituan; Yale University; Institute of Automation, Chinese Academy of Sciences(南京大学; 美团; 耶鲁大学; 中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型的安全评估,提出细粒度基准MME-Safety,采用四维标注和分层评估框架,在17个模型上揭示跨模态风险及思维链推理的安全隐患。

AI 中文摘要

尽管多模态大语言模型(MLLMs)展现出显著进步,但其跨模态能力引入了复杂的脆弱性,这些脆弱性容易绕过单模态过滤器。现有基准缺乏细粒度的意图相关标注,并依赖单一维度指标,阻碍了全面的鲁棒性评估。为解决这一问题,我们提出了MME-Safety,一个经过严格验证的基准,具有独特的四维标注模式,对风险场景、危害严重性和模态特定隐蔽级别进行分类。此外,我们引入了一个分层评估框架,以评估基础响应可靠性、实际风险暴露以及防御行为的结构完整性。对17个最先进MLLMs的广泛零样本评估提供了当前多模态系统的全面安全概况。我们的分析系统性地研究了跨模态输入配置,并揭示了与思维链(CoT)推理相关的安全隐患。这些多方面的发现强调了在多模态领域进行稳健的、具有推理意识的安全对齐的紧迫需求。

英文摘要

While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related annotations and rely on unidimensional metrics, hindering comprehensive robustness evaluation. To address this, we propose MME-Safety, a rigorously verified benchmark featuring a unique four-dimensional annotation schema that categorizes risk scenarios, harm severity, and modality-specific stealth levels. Furthermore, we introduce a hierarchical evaluation framework to assess fundamental response reliability, actual risk exposure, and the structural integrity of defensive behaviors. Extensive zero-shot evaluations across 17 state-of-the-art MLLMs provide a comprehensive safety profile of current multimodal systems. Our analysis systematically investigates cross-modal input configurations and uncovers safety implications associated with Chain-of-Thought (CoT) reasoning. These multifaceted findings underscore the urgent need for robust, reasoning-aware safety alignment in the multimodal landscape.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑