AI 中文总结
本研究针对现有语义偏移越狱攻击有效性有限的问题,提出带迭代上下文优化(ICO)的黑盒上下文感知框架,经多数据集与多模型实验,其攻击效果优于现有基线,平均成功率达74.6%。
AI 中文摘要
基础模型在各类任务中取得了显著成功,但仍存在脆弱性。为探究此类脆弱性,语义偏移越狱攻击近期成为颇具前景的攻击范式:该方法通过将原始有害问题中的有害术语替换为良性替代词,并利用上下文信息诱导目标模型将这些替代词重新解释为对应的有害概念,从而绕过明确的安全机制。然而,现有语义偏移越狱攻击的有效性往往有限。在本研究中,我们揭示该局限性源于对上下文语义偏移能力的忽视。通过系统分析,我们发现上下文在诱导语义偏移方面展现出显著不同的能力:具备更强语义偏移能力的上下文更易引导模型恢复有害含义,从而实现成功的越狱攻击。基于这一发现,我们系统识别并提炼有效上下文的特征,提出一种带迭代上下文优化(ICO)的黑盒上下文感知语义偏移越狱攻击框架。在每一轮迭代中,ICO利用这些特征及目标模型的反馈优化上下文。在三个数据集和八个目标基础模型上开展的大量实验表明,ICO始终优于八个最先进的基线方法,平均攻击成功率达74.6%。
英文摘要
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful terms in original harmful questions with benign alternatives and leveraging contextual information to induce the target model to reinterpret these alternatives as their corresponding harmful concepts. However, existing semantic-shift jailbreaks often achieve limited effectiveness. In this work, we reveal that this limitation arises from overlooking the semantic-shift capability of contexts. Through systematic analysis, we find that contexts exhibit substantially different abilities in inducing semantic shifts: contexts with stronger semantic-shift capabilities are more likely to guide models toward recovering harmful meanings and achieving successful jailbreaks. Based on this finding, we systematically identify and distill the characteristics of effective contexts and propose a black-box context-aware semantic-shift jailbreak framework with Iterative Context Optimization (ICO). In each iteration, ICO leverages these characteristics and feedback from the target model to optimize contexts. Extensive experiments on three datasets and eight target foundation models demonstrate that ICO consistently outperforms eight state-of-the-art baselines, achieving an average attack success rate of 74.6%.