arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨架构视觉语言模型中基于方向的推理时防御措施审计

A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models

Xiangyu Yin, Tora Bodin, Rohan Menon, Chih-Hong Cheng

arXiv 2607.27910首次发表:更新:

AI 中文总结

该研究审计了视觉语言模型中5种基于方向的推理时防御,发现其性能具架构特异性,部分防御间几何重叠,需按解码器家族单独校准。

AI 中文摘要

针对视觉语言模型越狱的推理时防御,通常会在选定的解码器层从残差流中减去校准后的方向。我们在幅度控制协议下,对来自四个架构家族的15个模型与层单元,比较了五个防御候选方案,该协议为每个提示匹配干预规模,并将每个方向与具有相同范数的随机对照配对。候选方案包括:平均图像条件偏移、CMRM风格的弃权(不执行)方向、ShiftDC风格的攻击特定残差、提示忽略图像的指令以及随机对照。没有单个候选方案在弃权(不执行)恢复和效用保留两方面均占优。图像条件偏移在LLaVA 1.5和Pixtral 12B中表现领先,且是唯一在所有家族中效用损失保持在测量噪声下限的候选方案。提示指令在Qwen2.5 VL中领先,而攻击特定残差在Qwen2 VL 2B中领先。图像条件方向在15个单元中的13个具有方向特异性,但具有强架构特异性,且在唯一兼容维度对(LLaVA 1.5 13B与Pixtral 12B)间不可迁移。我们还连接了纯文本与多模态弃权(不执行)几何:CMRM方向在所有15个单元中与图像条件偏移具有正余弦对齐,均值为0.35,范围0.17至0.65,是随机向量零的15至25倍,符号检验p值约为3e-5。这些结果表明,两种方案恢复了部分重叠的几何,且基于方向的防御应针对每个语言解码器家族单独校准。

英文摘要

Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder layer. We compare five defence candidates across 15 model and layer cells from four architectural families under a magnitude controlled protocol that matches the intervention size for each prompt and pairs every direction with a random control of the same norm. The candidates are the mean image conditioning shift, a CMRM style refusal direction, a ShiftDC style attack specific residual, a prompt instruction to ignore the image, and a random control. No single candidate dominates on both refusal recovery and utility preservation. The image conditioning shift leads on LLaVA 1.5 and Pixtral 12B and is the only candidate whose utility loss remains at the measurement noise floor in every family. The prompt instruction leads on Qwen2.5 VL, while the attack specific residual leads on Qwen2 VL 2B. The image conditioning direction is direction specific in 13 of 15 cells, but strongly architecture specific and nontransferable across the only dimension compatible pair, LLaVA 1.5 13B and Pixtral 12B. We also connect text only and multimodal refusal geometry. The CMRM direction has positive cosine alignment with the image conditioning shift in all 15 cells, with mean 0.35, range 0.17 to 0.65, 15 to 25 times the random vector null, and a sign test p value of about 3e-5. These results show that the two recipes recover partially overlapping geometry and that direction based defences should be calibrated separately for each language decoder family.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑