arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.03416cs.CV

GAMBIT:一种用于多模态大语言模型的游戏化突破框架

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models

  • Georgia State University(佐治亚州立大学)
  • Shandong University(山东大学)
  • Nanyang Technological University, Singapore(新加坡南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiangdong Hu, Yangyang Jiang, Qin Hu, Xiaojun Jia

更新

AI总结:

本文提出GAMBIT框架,通过分解重组有害视觉语义并构建游戏化场景,使模型在探索和重建意图中主动完成突破,提升攻击成功率。

AI中文摘要:

多模态大语言模型(MLLMs)已被广泛应用,但其在对抗输入下的安全对齐仍脆弱。先前研究显示增加推理步骤会破坏安全机制,导致MLLMs生成攻击者期望的有害内容。然而,大多数现有攻击仅增加修改视觉任务的复杂性,未显式利用模型自身推理激励。这导致其在推理模型(具有思维链的模型)上表现劣于非推理模型(无思维链的模型)。如果模型能像人类思考,能否影响其认知阶段决策使其主动完成突破?为验证此想法,我们提出GAMBI(通过指令陷阱进行游戏化对抗多模态突破),一种新的多模态突破框架,分解并重组有害视觉语义,然后构建游戏化场景,驱动模型探索、重建意图并回答以赢得游戏。所生成的结构化推理链在视觉和文本方面增加任务复杂性,使模型成为目标追求者,其目标追求减少安全关注并导致其回答重构的恶意查询。在流行的推理和非推理MLLMs上的大量实验表明,GAMBIT实现了高攻击成功率(ASR),在Gemini 2.5 Flash上达到92.13%,在QvQ-MAX上达到91.20%,在GPT-4o上达到85.87%,显著优于基线。

英文摘要:

Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous work has shown that increasing inference steps can disrupt safety mechanisms and lead MLLMs to generate attacker-desired harmful content. However, most existing attacks focus on increasing the complexity of the modified visual task itself and do not explicitly leverage the model's own reasoning incentives. This leads to them underperforming on reasoning models (Models with Chain-of-Thoughts) compared to non-reasoning ones (Models without Chain-of-Thoughts). If a model can think like a human, can we influence its cognitive-stage decisions so that it proactively completes a jailbreak? To validate this idea, we propose GAMBI} (Gamified Adversarial Multimodal Breakout via Instructional Traps), a novel multimodal jailbreak framework that decomposes and reassembles harmful visual semantics, then constructs a gamified scene that drives the model to explore, reconstruct intent, and answer as part of winning the game. The resulting structured reasoning chain increases task complexity in both vision and text, positioning the model as a participant whose goal pursuit reduces safety attention and induces it to answer the reconstructed malicious query. Extensive experiments on popular reasoning and non-reasoning MLLMs demonstrate that GAMBIT achieves high Attack Success Rates (ASR), reaching 92.13% on Gemini 2.5 Flash, 91.20% on QvQ-MAX, and 85.87% on GPT-4o, significantly outperforming baselines.

补充信息

↑