发表机构
Shanghai Jiao Tong University; University of Illinois Urbana-Champaign; Zhejiang University; Nanyang Technological University(上海交通大学; 伊利诺伊大学厄巴纳-香槟分校; 浙江大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出面向意图的多轮LLM越狱分类法,发现攻击有效性取决于意图组织精细度,推动检测范围升级,表明轮级安全机制不足,需对应层级评估协议。
AI 中文摘要
大语言模型(LLMs)越来越多地被部署在交互式场景中,用户意图通常通过多轮对话逐步展开。多轮越狱攻击利用这一模式,在多轮对话中逐步推进有害意图,使得没有单条消息暴露完整目标。然而,现有研究将这些攻击视为零散的提示模式集合,未分析攻击者如何在交互中组织和推进有害意图。我们提出了一个由四部分组成的、面向意图的分类法,根据攻击者的意图结构对多轮越狱攻击进行组织。通过受控消融实验,我们发现攻击有效性由意图在多轮中的组织精细度决定,而非上下文长度或查询数量。我们进一步表明,意图的组织方式决定了其可被检测的程度,将所需检测范围从轮级推至会话级,再到跨会话级。这些发现表明,轮级局部安全机制在结构上不足,单点评估忽略了意图的组织方式,推动了与有害意图可被观测的层级对齐的评估协议。代码可在以下网址获取:this https URL。
英文摘要
Large Language Models (LLMs) are increasingly deployed in interactive settings, where user intent commonly unfolds through multi-turn dialogue. Multi-turn jailbreaks exploit this pattern by advancing a harmful intent across turns, so that no single message exposes the full objective. However, existing work treats these attacks as a loose collection of prompt patterns and does not analyze how the adversary organizes and advances harmful intent across an interaction. We develop a four-part, intent-oriented taxonomy that organizes multi-turn jailbreaks by adversarial intent structure. Through controlled ablations, we find that effectiveness is driven by how deliberately intent is organized across turns rather than by context length or query count. We further show that the way intent is organized determines the level at which it becomes detectable, pushing the required detection surface outward from the turn level to the session level to the cross-session level. These findings indicate that turn-local safety mechanisms are structurally insufficient and that single-point evaluation overlooks how intent is organized, motivating evaluation protocols aligned to the level at which harmful intent becomes observable. The code is available at: https://github.com/SiyuanLi00/INTACT.