TempJail:基于字幕时序调度的针对大型视觉语言模型的时序越狱攻击
TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling
- University of Electronic Science and Technology of China(电子科技大学)
- East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TempJail是一种黑盒视频越狱框架,通过优化字幕时序调度攻击LVLMs,在四个模型和两个数据集上的攻击成功率优于基线。
AI中文摘要:
大型视觉语言模型(LVLMs)在视频理解与推理领域已取得显著进展。尽管针对文本和图像的越狱攻击已有大量研究,但针对LVLMs的视频越狱攻击仍大多未被探索。现有视频越狱方法主要操控嵌入视频中的文本内容,却忽略了此类信息的时间组织方式。我们的分析表明,越狱效果不仅取决于文本信息的语义,还取决于其时序呈现,包括时长及时隙分配。基于这一发现,我们采用现实视频中常见的字幕作为自然攻击媒介,字幕可在不造成视觉干扰的前提下,对语义内容进行精确时序控制。基于此,我们提出TempJail,这是一种基于视频的黑盒越狱框架,其构建与查询对齐的对话式字幕序列,并优化其时序调度,以利用LVLMs的时序漏洞,诱导出满足源查询有害意图的响应。在四个代表性LVLMs和两个数据集上开展的大量实验表明,TempJail在所有评估的模型-数据集设置中均达到最高攻击成功率(ASR),在GPT-5和Gemini 3.5-Flash上的数据集平均ASR分别比最强基线高出53和18个百分点。
英文摘要:
Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored. Existing video jailbreak methods mainly manipulate textual content embedded in videos, while overlooking how such information is organized over time. Our analysis reveals that jailbreak effectiveness depends not only on the semantics of textual information but also on its temporal presentation, including duration and timing-slot allocation. Motivated by this finding, we use subtitles, which are common in real-world videos and allow semantic content to be presented under precise temporal control without appearing visually intrusive, as a natural attack medium. Based on this insight, we propose TempJail, a black-box video-based jailbreak framework that constructs query-aligned dialogue-style subtitle sequences and optimizes their temporal scheduling to exploit temporal vulnerabilities in LVLMs and elicit responses that satisfy the harmful intent of the source query. Extensive experiments on four representative LVLMs and two datasets demonstrate that TempJail achieves the highest attack success rate across all evaluated model--dataset settings, outperforming the strongest baseline by 53 and 18 percentage points in dataset-averaged ASR on GPT-5 and Gemini 3.5-Flash, respectively.