AI 中文总结
该研究针对多模态安全护栏提出不安全诱导攻击,通过不安全语义蒸馏实现84%攻击成功率,揭示了当前多模态安全架构存在安全输入被误判为不安全的可用性漏洞。
AI 中文摘要
多模态安全护栏模型已成为视觉-语言系统中内容筛查的关键安全组件。尽管对抗性研究已广泛研究了产生漏报的越狱攻击,但诱导良性输入产生误报的反向威胁仍未被探索。我们提出了不安全诱导攻击,攻击者会分发经不可感知扰动的安全图像,触发安全护栏模型拒绝合法用户请求,造成“狼来了”效应,降低服务可用性并损害信任,这揭示了部署的安全过滤器中的可用性失效模式。为在不同用户提示下实现该威胁,我们提出了不安全语义蒸馏(USD),它将对抗性扰动与不安全内容的分布表示对齐,而非与特定提示实例对齐。在现实用户模拟场景中对四个最先进的安全护栏模型进行评估,USD实现了84%的攻击成功率,优于现有方法,暴露了当前多模态安全架构的基本漏洞。
英文摘要
Multimodal guard models have emerged as critical safety components for screening content in vision-language systems. While adversarial research has extensively studied jailbreaking attacks that produce false negatives, the inverse threat of inducing false positives on benign inputs remains unexplored. We introduce Unsafe Induction Attacks, where adversaries distribute imperceptibly perturbed safe images that trigger guard models to reject legitimate user requests, causing a "Boy Who Cried Wolf" effect that degrades service availability and erodes trust. This reveals an availability failure mode in deployed safety filters. To realize this threat under diverse user prompts, we propose Unsafe Semantic Distillation (USD), which aligns adversarial perturbations with distributional representations of unsafe content rather than prompt-specific instances. Evaluated on four state-of-the-art guard models across realistic user simulation scenarios, USD achieves 84% attack success rates, outperforming existing methods and exposing fundamental vulnerabilities in current multimodal safety architectures.
CommentsAccepted by KDD 2026