arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12373cs.AI

不想让你的大语言模型(LLM)推荐核打击?试试用日语提问

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

Rian Touchent

中文总结 AI 辅助

该研究发现,用日语提问可降低部分LLM在核打击场景中的发动率,其机制为模型的推理语言而非输入语言,且仅适用于在英语中已存在犹豫的模型,表明LLM安全行为具语言依赖性。

中文摘要 AI 辅助

大型语言模型(LLM)正越来越多地被用于战略和咨询场景,但其安全对齐通常仅在英语中进行评估。我们测试了来自6家提供商的9种模型,探究提示的语言是否能在高风险场景中改变模型的决策。我们采用单轮博弈论小插曲,让模型为一个拥有核武器的国家提供是否对无防御能力的对手发动打击的建议,提示在不同语言中均保持非道德性且战略上完全相同。我们发现,日语提示会降低Claude模型家族的发动打击率:在打击不必要的场景中,Claude Sonnet 4.6的发动率从40%降至0%;在有争议的场景中,从93%降至17%;而当打击在战略上合理时,影响极小。该效果也延伸到Gemini Pro 3.1(发动率从53%降至13%)。一项跨语言实验分离出了机制:当在英语提示中要求用日语推理时,发动率从93%降至37%。驱动该效果的是模型被要求进行推理的语言,而非输入的语言。当用日语推理时,模型会自发生成提示中完全未出现的道德词汇(如“道德成本”“数百万生命”)。其他5种模型未显示出语言效应,但无论语言如何,它们在几乎所有条件下都会发动打击。该效应需要模型在英语中已存在犹豫。这些结果表明,LLM的安全行为具有语言依赖性,仅用英语评估会遗漏其他语言中编码的风险和安全保障。

英文摘要

Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages.

补充信息

↑