发表机构
California Institute of Technology(加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多语言大语言模型安全对齐问题,引入涵盖多种语言、资源层级和扰动类型的Minionese基准测试及几何机理分析,发现不同攻击类型漏洞特征不同,指出仅英语安全评估不足,需考虑多因素。
AI 中文摘要
大语言模型中的安全对齐在跨语言环境中仍然很脆弱:在英语环境中被可靠拒绝的提示在非英语和低资源环境中可能会引发有害的合规行为。我们引入了Minionese,这是一个涵盖18种语言、4种资源层级和4种扰动类型(标准翻译、代码切换、音译和翻译腔)的多语言越狱基准测试,并对跨语言层级的拒绝失败进行了几何机理分析。我们表明,每种攻击类型都会产生不同的漏洞特征:音译漏洞由脚本身份介导,代码切换在最低资源层级保持有效性,所有模型在第2层和第3层之间都存在明显的安全状态转变。从机理上讲,低资源越狱通过将有害内容路由到几何上未对齐的子空间来成功,该子空间在拒绝方向上的投影不足,使拒绝机制保持完整但未触发。这些发现表明,仅英语的安全评估是不够的;它们需要考虑脚本家族、扰动类型和每种语言的对齐覆盖范围。基准测试和分析代码可在这个https网址获取。
英文摘要
Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each attack type produces a distinct vulnerability profile: transliteration vulnerability is mediated by script identity, code-switching maintains effectiveness through the lowest-resource tier, and a sharp safety regime transition between Tiers 2 and 3 is consistent across all models. Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometrically misaligned subspace that projects insufficiently onto the refusal directions, leaving the refusal mechanism intact but untriggered. These findings show that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage. The benchmark and analysis code is at https://github.com/Brentkong/Minionese-Comprehensive-Benchmark-and-Mechanistic-Study-of-Multilingual-LLM-Safety.git.