发表机构
Vrije Universiteit Brussel; imec-SMIT; Harvard University(布鲁塞尔自由大学; imec-SMIT(比利时微电子研究中心-可持续移动与信息技术研究组); 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究Qwen和Llama系列小权重模型中受扰动越狱输入的内部表示,选两个表示空间,在拒绝主导答案集中两空间均无行为超平面,仅特定模型的个别token与合规答案有关联。
AI 中文摘要
将不成功的越狱提示转变为成功提示的扰动技术不断演变,对大语言模型(LLM)安全构成重大威胁。本文研究了Qwen-2.5-1.5B/-3B/-7B-Instruct和Llama-3.2-1B/-3B/-3.1-8B-Instruct系列小权重模型中此类字符串级受扰动越狱输入的内部表示。选择了两个表示空间:最后一层最后一个token嵌入空间和前50个下一个token概率空间。在拒绝主导的答案集中,两个空间均未发现行为超平面。只有1.5B Qwen模型中的下一个token“Sure”,以及1$ Llama模型中的“,”和“ĊĊ”与合规标记答案有显著关联。
英文摘要
Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and "ĊĊ" in the 1$ Llama model, display a significant association with a compliant-labeled answer.
Comments21 pages, 9 figures, 7 tables, 2nd Workshop on Safe AI (SafeAI)