arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SceneJail:利用视频场景上下文越狱多模态大语言模型

SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs

Wenyu Chen, Li Wang, Chuanchao Zang, Xiangtao Meng, Xinyu Gao, Jianing Wang, Zheng Li, Shanqing Guo

arXiv 2609.38899首次发表:更新:

发表机构

Shandong University(山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SceneJail利用视频场景上下文作为攻击面,通过自适应场景构建和场景感知提示搜索,在黑盒条件下越狱视频多模态大语言模型,最高攻击成功率91.5%。

AI 中文摘要

视频多模态大语言模型(Video-MLLMs)支持对视频输入进行推理,但仍然容易受到越狱攻击,从而引发违反政策的响应。现有的视频越狱方法主要操纵有害查询的视觉呈现方式,因此仅将视频视为载体。因此,周围的视频场景作为上下文攻击面尚未被探索。在本文中,我们表明相同的有害查询在不同视频场景中会引发不同的安全响应。为了系统地利用这一漏洞,我们提出了SceneJail,一种自适应黑盒越狱框架,包含两个协调组件。自适应场景构建动态搜索与有害查询上下文兼容的周围场景。场景感知提示搜索利用黑盒响应反馈来搜索针对所选场景定制的文本指导。在HADES和SafeBench数据集上对八个Video-MLLMs(包括两个专有模型GPT-4.1和Gemini3.5-Flash)的广泛评估证明了SceneJail的有效性。SceneJail-F持续呈现完整查询,实现了高达91.5%的平均攻击成功率(ASR),比最强基线高出29.1个百分点。此外,SceneJail-S将查询分布在连续帧中,对当前防御保持高度鲁棒性,即使在严格的图像过滤下仍保留72.3%的ASR。

英文摘要

Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently, the surrounding video scenario remains unexplored as a contextual attack surface. In this paper, we show that the same harmful query can elicit different safety responses when placed in different video scenarios. To systematically exploit this vulnerability, we propose SceneJail, an adaptive black-box jailbreak framework with two coordinated components. Adaptive Scenario Construction dynamically searches for a surrounding scenario that is contextually compatible with the harmful query. Scenario-aware Prompt Search uses black-box response feedback to search for textual guidance tailored to the selected scenario. Extensive evaluations on the HADES and SafeBench datasets across eight Video-MLLMs, including two proprietary models, GPT-4.1 and Gemini3.5-Flash, demonstrate the effectiveness of SceneJail. SceneJail-F, which presents the complete query persistently, achieves average attack success rates (ASR) up to 91.5%, outperforming the strongest baselines by 29.1 percentage points. Furthermore, SceneJail-S, which distributes the query across successive frames, remains highly robust against current defenses, retaining a 72.3% ASR even under strict image filtering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑