SteerProbe:学习绕过视觉-语言模型中的安全转向
SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models
- The Hong Kong Polytechnic University(香港理工大学)
- Anhui University of Technology(安徽工业大学)
- Nanyang Technological University(南洋理工大学)
- Nanjing University of Science and Technology(南京理工大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
SteerProbe是一种仅输出的黑盒攻击,通过学习选择有效改写形式,在三个VLM骨干、两个基准和三种转向防御下,将有害率从7.43%提升至18.36%,揭示转向防御在改写输入下的脆弱性。
中文摘要 AI 辅助
激活转向通过修改中间表示而不更新骨干网络参数,为视觉-语言模型(VLM)提供了一种推理时的防御手段。然而,针对基准输入的保护可能无法在相同有害请求的替代表达形式下持续生效。我们利用旨在保留潜在意图的固定文本、视觉和联合改写形式来研究这一差距,并发现这些更改可以绕过具有代表性的转向防御。一项补充性的局部分析提供了一个充分条件,在该条件下,尽管局部转向修正发生任何可允许的变化,改写形式仍能跨越替代安全边界。随后,我们引入了SteerProbe,一种仅基于输出的黑盒攻击,它从共享的校准预算中学习为未见过的请求选择有效的改写形式。在三个VLM骨干网络、两个基准和三种转向防御的设置下,SteerProbe在每个端点和基准使用总计500次校准查询,在所有受防御的设置中均提高了有害率,将平均值从7.43%提升至18.36%。这些发现强调,原始基准输入上的鲁棒性不足以表征转向防御的安全性,并促使将改写鲁棒性作为重要的评估维度。它们进一步推动了在保持良性效用的同时,跨意图保留的多模态变体保持安全性的转向机制的发展。
英文摘要
Activation steering offers an inference-time defense for vision--language models (VLMs) by modifying intermediate representations without updating backbone parameters. However, protection on benchmark inputs may not persist across alternative expressions of the same harmful request. We investigate this gap using fixed textual, visual, and joint reformulations designed to preserve the underlying intent, and find that these changes can bypass representative steering defenses. A complementary local analysis provides a sufficient condition under which a reformulation can cross a surrogate safety margin despite any admissible change in the local steering correction. We then introduce SteerProbe, an output-only black-box attack that learns to select effective reformulations for unseen requests from a shared calibration budget. Across three VLM backbones, two benchmarks, and three steering defenses, SteerProbe increases Harmful Rate in all defended settings using 500 total calibration queries per endpoint and benchmark, raising the average from 7.43% to 18.36%. These findings highlight that robustness on original benchmark inputs is insufficient to characterize the safety of steering defenses and motivate reformulation robustness as an important evaluation dimension. They further motivate steering mechanisms that preserve safety across intent-preserving multimodal variations while maintaining benign utility.