发表机构
AI4S Center, Shanghai Qi Zhi Institute; College of AI, Tsinghua University; Fangcun AI; Institute for Interdisciplinary Information Sciences, Tsinghua University(上海期智研究院AI4S中心; 清华大学人工智能学院; 方寸人工智能; 清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出BLUEPRINT安全评估框架,结合WORLDVIEWSIM模块与蒙特卡洛树搜索优化18种心理学影响因子,在6个前沿模型上实现接近上限的攻击成功率,揭示多轮越狱的漏洞机制并为安全评估提供新方向。
AI 中文摘要
多轮越狱攻击表明,有害意图可分布在对话过程中,但现有方法未明确驱动漏洞的对话机制。我们提出BLUEPRINT(一种安全评估框架),它将因子化的社会影响策略空间与WORLDVIEWSIM(跨轮情境上下文模块)分离。蒙特卡洛树搜索优化了4轮对话轨迹中18种基于理论的影响因子的轮级组合。在6个前沿模型上,BLUEPRINT在主要开放权重模型和专有模型上达到接近上限的攻击成功率(ASR),且所需平均查询数最少(2.46)。生成的轨迹进一步揭示了抵抗目标中的模型特定漏洞:每个模型对不同的影响因子和策略转换有不同反应,但都有共同的恢复路径——转向具体、可执行的任务框架始终能摆脱硬拒绝状态。消融实验证实操作线索最重要:使请求可操作影响最大,收益框架异常有效,部分合法性诉求可能产生反效果。这些发现表明,稳健的多轮安全不仅需要监测有害内容,还需监测对话状态如何使不安全请求显得具体且可本地执行。
英文摘要
Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier models, BLUEPRINT achieves near-ceiling ASR on major open-weight and proprietary models, while requiring the fewest average queries (2.46). The resulting trajectories further reveal model-specific vulnerability among resistant targets: each responds to distinct influence factors and strategy transitions, yet all share a common recovery pathway-shifting toward concrete, executable task framing consistently escapes hard-refusal states. Ablations confirm operational cues matter most: making requests actionable has the largest impact, gain framing is unusually potent, and some legitimacy appeals can backfire. These findings suggest robust multi-turn safety requires monitoring not only harmful content, but also how dialogue state makes unsafe requests appear concrete and locally executable.
Comments19 pages, 7 figures. Accepted to Findings of EMNLP 2026