CodeMimicry:通过结构化代码补全利用大型语言模型中的安全泛化滞后
CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
浏览论文内容
中文总结 AI 辅助
针对大模型安全对齐在代码领域泛化滞后的问题,提出CodeMimicry黑盒越狱框架,通过结构化代码补全实现96.25%攻击成功率,并揭示其绕过安全机制的机理。
中文摘要 AI 辅助
大型语言模型在多个领域展现出了卓越的能力,但其安全对齐仍然容易受到越狱攻击。在本工作中,我们识别出一个此前未被充分探索的失败模式——安全泛化滞后,即主要基于自然语言训练的对齐未能迁移到代码领域。我们表明,这种滞后导致了代码补全盲区,使得嵌入在语法有效代码中的恶意意图能够规避安全机制。为了利用这一漏洞,我们提出了CodeMimicry,一个全自动的黑盒越狱框架,通过生成结构化的面向对象代码提示来诱导代码补全产生有害输出。在8个最先进的商业大语言模型上的实验表明,CodeMimicry实现了96.25%的攻击成功率,平均仅需1.51次查询,显著优于基于模板和基于优化的基线方法。除了实证性能之外,我们还通过潜在空间表征对基于代码的越狱进行了机制分析,包括投影到拒绝相关方向以及激活引导。该分析解释了CodeMimicry如何在代码相关领域绕过安全机制。我们的发现揭示了当前安全对齐中的弱点,并强调了在代码等结构化领域中实现稳健对齐的必要性。
英文摘要
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
发表机构
- Zhejiang Sci-Tech University(浙江理工大学)
- China Academy of Information and Communications Technology(中国信息通信研究院)
机构由 AI 辅助整理,请以论文原文为准。