发表机构
Renmin University of China; Xi’an Jiaotong University; University of Science and Technology of China; Wuhan University; Shanghai Artificial Intelligence Laboratory(中国人民大学; 西安交通大学; 中国科学技术大学; 武汉大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型在特定场景下的脆弱性,提出Concept2Scenario框架,通过实例化概念空间、归因拒绝抑制等步骤发现脆弱场景,在多模型和基准测试中验证,能提高攻击成功率、可迁移且组合效果优。
AI 中文摘要
安全对齐的大语言模型被训练来拒绝有害请求,但将相同请求嵌入特定场景可能绕过其防护机制。现有红队测试方法通过观察攻击结果凭经验识别有效场景,但特定场景削弱拒绝的机制尚不清楚。同时,机制可解释性研究刻画了拒绝方向和越狱相关特征,但未解释两者关系。本文表明场景包装提示激活内部场景方向,其因果导向持续降低拒绝分数。在此基础上,提出基于概念的归因框架Concept2Scenario用于发现脆弱场景,通过稀疏自动编码器实例化广泛概念空间,将拒绝抑制归因于单个概念,转化为可解释的自然语言场景,并通过交互归因识别协同场景组合。在三个开源模型、两个安全基准和六种黑盒越狱方法上,发现的场景可提高平均攻击成功率达18.2个百分点,还可迁移到其他模型,且识别出的组合优于单个成分,能使迭代攻击在更少轮次成功。
英文摘要
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relationship between the two representations. In this work, we show that scenario-wrapped prompts activate internal scenario directions whose causal steering consistently reduces refusal scores. Building on this finding, we propose \textsc{Concept2Scenario}, a concept-based attribution framework for vulnerable scenario discovery. It instantiates a broad concept space with a sparse autoencoder, attributes refusal suppression to individual concepts, translates the identified concepts into interpretable natural-language scenarios, and identifies synergistic scenario combinations through interaction attribution. Across three open-source models, two safety benchmarks, and six black-box jailbreak methods, the discovered scenarios serve as reusable priors that improve average attack success rates by up to $18.2$ percentage points. They also transfer to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, suggesting that some scenario-level refusal vulnerabilities are shared across model families. Moreover, the identified combinations outperform their individual constituents and enable iterative attacks to succeed in fewer turns.
Comments19 pages, 11 Figures, Under Review