发表机构
New York University Abu Dhabi; Tandon School of Engineering, New York University(纽约大学阿布扎比分校; 纽约大学坦登工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究剖析图像到文本越狱中图像侧因素,发现攻击图像显著提升成功率,仅相关性操纵(E4)有可归因效应,其余因素作用有限。
AI 中文摘要
图像到文本的越狱将有害意图置于文本、图像内容或它们之间的关系中。我们在一个包含313个提示的StrongREJECT切片上,跨五个多模态模型,并额外在附录中对InternVL3.5-8B进行评估,检查了四个已发表攻击家族中的图像侧因素。有害指令在所有条件下保持不变;基线矩阵每个提示使用一次抽取,配对消融使用三次抽取并采用自动化评分标准评判。一个裸露的有害查询,无论是否附带良性无关图像,在大多数受害者上产生的攻击成功率都很低,而攻击图像则显著提高了成功率。首先,逐块熵和JPEG大小无法可靠地区分攻击块与大小匹配的良性干扰块,限制了仅基于密度的筛选。其次,早期的块数阶梯被有效载荷可见性所混淆。一个修正的区域计数测试未发现可检测的影响,因此块数结构的作用仍未解决。第三,在Qwen3-VL-8B上,移除查询特定相关性的E4操纵将ASR降低了约0.12。这支持了对相关性操纵的有界归因,尽管图像-文本一致性仍未测量。一个类别内对照在重叠刺激上重现了该方向。五个测试通过了全局统计校正,但只有E4支持对单个测量描述符的归因;类别内结果是稳健性检查,而非独立归因。这些结论仍以评分标准评判为条件。
英文摘要
Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The harmful instruction is held constant across conditions; the baseline matrix uses one draw per prompt, and paired ablations use three draws with an automated rubric judge. A bare harmful query, with or without a benign unrelated image, produces little attack success on most victims, while attack images substantially increase it. First, per-tile entropy and JPEG size do not reliably distinguish attack tiles from size-matched benign distractors, limiting density-only screening. Second, earlier tile-count ladders were confounded by payload visibility. A corrected region-count test found no detectable effect, so the role of tile-count structure remains unresolved. Third, on Qwen3-VL-8B, the E4 manipulation that removes query-specific relatedness lowers ASR by about 0.12. This supports a bounded attribution to the relatedness manipulation, although image-text congruence remains unmeasured. A within-category control reproduces the direction on overlapping stimuli. Five tests survive the global statistical correction, but only E4 supports attribution to one measured descriptor; the within-category result is a robustness check, not a separate attribution. These conclusions remain conditional on the rubric judge.