安全提醒:一种软提示以重新激活视觉语言模型中延迟的安全意识
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
- School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
- School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院)
- School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算与信息系统学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视觉语言模型存在延迟安全意识的现象,提出安全提醒软提示调优方法,通过优化可学习提示令牌定期注入以增强安全意识,显著降低攻击成功率并保持模型性能。
AI中文摘要:
随着视觉语言模型(VLMs)在代码生成和聊天助手等实际应用中展现出日益强大的能力,确保其安全性变得至关重要。与传统的大型语言模型(LLMs)不同,视觉语言模型因其多模态特性而面临独特的脆弱性,攻击者可以修改视觉或文本输入以绕过安全护栏并触发有害内容的生成。通过对攻击下视觉语言模型行为的系统分析,我们识别出一种称为“延迟安全意识”的新现象。具体而言,我们观察到经过安全对齐的视觉语言模型可能最初被攻破以产生有害内容,但最终会识别相关风险并尝试自我纠正。这一模式表明,视觉语言模型保留了其潜在的安全意识,但其激活存在时间延迟。基于这一见解,我们假设可以通过精心设计的提示主动重新激活视觉语言模型的安全意识。为此,我们引入了“安全提醒”,一种软提示调优方法,优化可学习的提示令牌,在文本生成过程中定期注入以增强安全意识,有效防止有害内容的生成。此外,我们的安全提醒仅在检测到有害内容时激活,不影响正常对话,并保持模型在良性任务上的性能。通过在三个既定安全基准和一个对抗性攻击上的全面评估,我们证明我们的方法显著降低了攻击成功率,同时保持了模型效用,为在实际应用中部署更安全的视觉语言模型提供了一种实用解决方案。
英文摘要:
As Vision-Language Models (VLMs) demonstrate increasing capabilities across real-world applications such as code generation and chatbot assistance, ensuring their safety has become paramount. Unlike traditional Large Language Models (LLMs), VLMs face unique vulnerabilities due to their multimodal nature, allowing adversaries to modify visual or textual inputs to bypass safety guardrails and trigger the generation of harmful content. Through systematic analysis of VLM behavior under attack, we identify a novel phenomenon termed ``delayed safety awareness''. Specifically, we observe that safety-aligned VLMs may initially be compromised to produce harmful content, but eventually recognize the associated risks and attempt to self-correct. This pattern suggests that VLMs retain their underlying safety awareness but experience a temporal delay in their activation. Building on this insight, we hypothesize that VLMs' safety awareness can be proactively reactivated through carefully designed prompts. To this end, we introduce ``The Safety Reminder'', a soft prompt tuning approach that optimizes learnable prompt tokens, which are periodically injected during the text generation process to enhance safety awareness, effectively preventing harmful content generation. Additionally, our safety reminder only activates when harmful content is detected, leaving normal conversations unaffected and preserving the model's performance on benign tasks. Through comprehensive evaluation across three established safety benchmarks and one adversarial attacks, we demonstrate that our approach significantly reduces attack success rates while maintaining model utility, offering a practical solution for deploying safer VLMs in real-world applications.