针对大型语言模型的任意密码攻击无需微调
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
浏览论文内容
中文总结 AI 辅助
本文发现前沿大模型无需微调,仅通过提示和上下文学习即可掌握任意密码通信技能,从而绕过对齐和有害内容分类器,实现对Anthropic、Google和OpenAI模型的越狱攻击。
中文摘要 AI 辅助
大型语言模型的安全与安保研究,除其他事项外,主要关注检测和预防越狱攻击:即对齐绕过,允许对抗性用户从模型中引出不需要或有害的输出。任意密码(即隐蔽通信)攻击是此类越狱攻击的一种,此前已在商业模型的微调API上得到验证。在这些攻击中,目标模型在包含加密的有害问题及回答的语料库上进行训练,随后通过学到的加密方案对有害请求作出响应。在本文中,我们表明,更新的前沿模型无需微调即可获得基于密码的通信技能。相反,它们可以通过提示(prompting)并在必要时通过上下文学习(in-context learning)来学习这些技能。此外,当通过学到的密码进行通信时,模型的对齐能力显著减弱或被完全绕过。据我们所知,这构成了针对商业黑盒大型语言模型的一种新型攻击向量。我们成功演示了针对由Anthropic、Google和OpenAI开发的前沿模型的越狱攻击。我们的攻击绕过了商业有害内容分类器,因为有害内容被加密,因此看起来像无意义的文本或乱码。
英文摘要
Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arbitrary cipher, or covert communication, attacks are one such type of jailbreak and have previously been demonstrated against the fine-tuning APIs of commercial models. In these attacks, target models are trained on a corpus of encrypted harmful questions and responses and subsequently respond to harmful requests through the learned encryption scheme. In this paper, we show that newer frontier models do not require fine-tuning to acquire cipher-based communication skills. Instead, they can learn these skills through prompting and, when necessary, through in-context learning. Furthermore, model alignment is significantly weakened or entirely bypassed when communication occurs through the learned cipher. To the best of our knowledge, this constitutes a novel attack vector against commercial black-box large language models. We demonstrate successful jailbreaks against frontier models developed by Anthropic, Google, and OpenAI. Our attack bypasses commercial harmfulness classifiers because harmful content is encrypted and therefore appears as nonsensical text or gibberish.
发表机构
- McGill University(麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。