发表机构
Michigan Technological University; Miami University; Kansas State University(密歇根理工大学; 迈阿密大学; 堪萨斯州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型越狱攻击,提出基于深度提示优化的动态防御方法DDPO,利用模型中间层生成防御嵌入,无需修改权重,在多种模型和攻击上显著优于静态方法。
AI 中文摘要
大语言模型(LLMs)在许多应用中展现出令人印象深刻的能力,但仍然容易受到越狱攻击,这些攻击会诱导模型生成有害或非预期的内容。虽然模型微调是安全对齐的一种选择,但成本高昂且容易发生灾难性遗忘。提示优化已成为一种有前景的替代方案,然而现有的基于提示的防御通常依赖于静态修改(例如固定的前缀或后缀),无法适应多样且不断演变的攻击。我们提出了动态深度提示优化(DDPO),这是首个基于深度提示优化的越狱防御方法。DDPO利用目标大语言模型自身的中间层作为特征提取器,通过一个轻量级多层感知机动态生成防御性嵌入。这些定制化的嵌入随后被注入到后续的中间层,从而在不修改大语言模型权重的情况下实现依赖于输入的防御。这种设计确保了高适应性,同时计算开销极小。在多种模型和攻击上的实验表明,DDPO显著优于静态提示优化方法,尤其是在弱对齐模型以及处理语义模糊的良性提示时,能够成功地将它们与真正有害的请求区分开来。
英文摘要
Large Language Models (LLMs) demonstrate impressive capabilities across many applications but remain vulnerable to jailbreak attacks, which elicit harmful or unintended content. While model fine-tuning is an option for safety alignment, it is costly and prone to catastrophic forgetting. Prompt optimization has emerged as a promising alternative, yet existing prompt-based defenses typically rely on static modifications (e.g., fixed prefixes or suffixes) that cannot adapt to diverse and evolving attacks. We propose Dynamic Deep Prompt Optimization (DDPO), the first jailbreak defense based on deep prompt optimization. DDPO uses the target LLM's own intermediate layers as feature extractors to dynamically generate defensive embeddings via a lightweight multilayer perceptron. These tailored embeddings are then injected into a subsequent intermediate layer, enabling an input-dependent defense without modifying the LLM's weights. This design ensures high adaptability with minimal computational overhead. Experiments on a diverse set of models and attacks demonstrate that DDPO significantly outperforms static prompt optimization methods, particularly on weakly aligned models and when handling semantically ambiguous benign prompts, successfully distinguishing them from genuinely harmful requests.