arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构性越狱具有泛化性但不具有叠加性:跨提供商和多语言的非自愿上下文学习研究

Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning

Tejasvi C. Addagada

arXiv 2609.08373首次发表:更新:

AI 中文总结

本研究测试结构性越狱(IICL)与多语言安全弱化是否叠加,发现IICL跨提供商泛化且金融攻击更严重,但非英语条件反而削弱攻击,表明越狱漏洞不可叠加,主要风险是英语结构性攻击。

AI 中文摘要

对齐的语言模型在两种独立压力下失效:最近被形式化为非自愿上下文学习(IICL)的结构性越狱类别,它将有害请求重新表述为数据标注任务的最后一个缺失单元格,通过模式而非内容判断来完成;以及英语之外安全对齐的削弱。一个自然的假设是这些压力会叠加。我们直接测试了这一假设。使用确定性的IICL操作符和StrongREJECT风格的评分判断器,我们对两个Google Gemini模型在两个基准上进行了红队测试:来自HarmBench的30个一般危害行为和来自FinProof的30个金融滥用行为,每个行为在单次基线设置和四种语言(英语、西班牙语、印地语、阿拉伯语)的IICL设置下进行。首先,IICL泛化到第二个提供商,并且在金融领域更严重:它将攻击成功率从HarmBench上的≤6.7%提升到80-90%,在FinProof上提升到97-100%,比其引入研究在OpenAI的GPT-5.4上报告的≤24%高出一个数量级。其次,与假设相反,将IICL输出强制转换为非英语语言并不会叠加两种弱点,反而会削弱攻击。十二个非英语条件中有十一个得分低于其英语基线(符号检验,p~0.003),唯一的例外是接近100%的天花板平局;在更强模型的金融集上,阿拉伯语从100%崩溃到33%。我们将此归因于相关性诅咒:一旦结构解锁了顺从性,模型在低资源语言中生成的有害内容质量较低,而实质评分判断器将其评为部分。该模式在独立的非Google判断器下重复(Cohen's kappa=0.86,377对配对判断),并且76.6%的非英语响应经过语言内验证。因此,越狱漏洞不是可加的;主导的残余风险是英语结构性攻击,对金融滥用最为严重,而不是多语言攻击。

英文摘要

Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety alignment outside English. A natural hypothesis is that these compound. We test it directly. Using a deterministic IICL operator and a StrongREJECT-style rubric judge, we red-team two Google Gemini models on two benchmarks, a 30 general-harm behaviours from HarmBench and 30 financial-abuse behaviours from FinProof, each under a single-shot baseline and under IICL in four languages (English, Spanish, Hindi, Arabic). First, IICL generalizes to a second provider and is worse in finance: it lifts attack success from <=6.7% to 80-90% on HarmBench and 97-100% on FinProof, an order of magnitude above the <=24% its introducing study reported on OpenAI's GPT-5.4. Second, against the hypothesis, forcing the IICL output into a non-English language does not stack the two weaknesses, it attenuates the attack. Eleven of twelve non-English conditions score below their English baseline (sign test, p~0.003), the lone exception a ceiling tie near 100%; on the stronger model's financial set Arabic collapses from 100% to 33%. We attribute this to a relevance curse: once structure has unlocked compliance, the models produce lower-quality harmful content in lower-resource languages, which a substance-grading judge scores as partial. The pattern replicates under an independent non-Google judge (Cohen's kappa=0.86, 377 paired verdicts), and 76.6% of non-English responses were verified in-language. Jailbreak vulnerabilities are therefore not additive; the dominant residual risk is the English structural attack, most acute for financial abuse, not a multilingual one.

Comments6 pages, 2 figures, 1 table. Pilot study. Includes a cross-family judge-agreement check (kappa=0.86) and output-language verification

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑