发表机构
Kennesaw State University(肯尼索州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多模态大语言模型的参考依赖型安全缺陷,提出COMIC参考感知安全门控,通过解析操作-目标对评估安全性,提升了模型鲁棒性且保留良性效用。
AI 中文摘要
多模态大语言模型(MLLMs)正越来越多地被用于与截图、扫描文档、图表及其他视觉接地输入进行交互。这一转变带来了新的安全风险:在许多多模态越狱攻击中,提示词和图像单独来看均无害,仅当模型将看似良性的操作(如总结、翻译或遵循指令)与某个局部视觉目标绑定后,才会产生不安全行为。这暴露了当前多模态防御的结构性缺陷:现有防御大多将提示词-图像对作为整体进行 moderation,而真正与安全相关的单元是解引用过程中产生的接地操作-目标对。在本研究中,我们识别并分析了这种依赖参考的失败模式,证明当有害语义被局部化、仅在接地后激活且依赖视觉参考解析时,现有防御会失效。为解决该问题,我们提出了COMIC(Context-Operation-Modality-Image-Classifier),一种面向MLLMs的预生成参考感知安全门控。COMIC首先推断请求的操作和参考类型,从OCR(光学字符识别)和开放词汇提案中构建候选目标,接地合理的指称对象,并针对明确的操作-目标对评估安全性。为保守处理歧义,COMIC在决定转发或阻止请求前,将最大风险聚合与质量感知路由相结合。我们在多个开源MLLMs、局部及更广泛的多模态越狱基准、以及良性参考敏感设置中对COMIC进行了评估。结果显示,COMIC在持续提升鲁棒性的同时,保留了良性效用和实际效率。更广泛而言,我们的发现表明,若不对请求的操作、其应用的视觉目标及该接地的置信度进行建模,就无法可靠地实施多模态安全。
英文摘要
Multimodal large language models (MLLMs) are increasingly used to interact with screenshots, scanned documents, diagrams, and other visually grounded inputs. This shift introduces a new safety risk: in many multimodal jailbreaks, neither the prompt nor the image is harmful in isolation. Unsafe behavior emerges only when the model binds an apparently benign operation, such as summarizing, translating, or following, to a localized visual target. This reveals a structural weakness in current multimodal defenses, which largely moderate the prompt-image pair as a whole even though the true security-relevant unit is the grounded operation-target pair produced during dereference. In this work, we identify and analyze this reference-dependent failure mode and show that existing defenses degrade when harmful semantics are localized, activated only after grounding, and dependent on visual reference resolution. To address this problem, we propose COMIC (Context-Operation-Modality-Image-Classifier), a reference-aware pre-generation safety gate for MLLMs. COMIC first infers the requested operation and reference type, constructs candidate targets from OCR and open-vocabulary proposals, grounds plausible referents, and evaluates safety over explicit operation-target pairs. To handle ambiguity conservatively, COMIC combines max-risk aggregation with quality-aware routing before deciding whether to forward or block a request. We evaluate COMIC across multiple open-source MLLMs, localized and broader multimodal jailbreak benchmarks, and benign reference-sensitive settings. The results show that COMIC consistently improves robustness while preserving benign utility and practical efficiency. More broadly, our findings suggest that multimodal safety cannot be enforced reliably without modeling the requested operation, the visual target to which it applies, and the confidence of that grounding.