目标对齐是否意味着目标恢复?对比编码器对抗性主张的证据阶梯研究
Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders
浏览论文内容
中文总结 AI 辅助
该研究探究视觉-语言模型的目标对齐(VTS)是否能预测独立模型的目标恢复,发现鲁棒训练的对比编码器传递更多独立证据,VTS仅在固定鲁棒编码器内有信息性,未发现跨编码器的单调对齐-证据关系。
中文摘要 AI 辅助
针对视觉-语言模型的对抗攻击会优化图像以匹配文本目标,随后将被攻击模型的相似度得分作为成功证据。我们探究该得分——受害者空间目标对齐(VTS)——是否能预测独立模型对目标的恢复。我们首先构建测量工具:被攻击几何外的监督评判者、真实目标混合对照、打乱目标负样本,以及从50%目标-图像混合得到的参考水平。随后两项预注册研究在匹配攻击下、三个扰动预算下对比六种对比编码器。经鲁棒训练的编码器(FARE、TeCoA、PMG、TRADES)比普通CLIP或SigLIP传递了多得多的独立证据;所有八项对比均在自举下限处被拒绝。然而,没有一个单元达到混合衍生的参考水平。三个最佳单元处于其复制带内,使得实际恢复情况未决。在鲁棒编码器内,单样本对齐增益与证据增益相关(ρ=0.24-0.51);在普通CLIP内,该相关性与零一致。跨编码器来看,我们未发现单调的对齐-证据关系。因此,VTS仅在固定鲁棒编码器内具有信息性,我们提供了替代它的报告协议。
英文摘要
Adversarial attacks on vision-language models optimize an image toward a text target, then cite the attacked model's similarity score as evidence of success. We ask whether that score - victim-space target alignment (VTS) - predicts recovery of the target by an independent model. We first build a measurement instrument: supervised judges outside the attacked geometry, real-target blend controls, shuffled-target negatives, and a reference level derived from a 50% target-image blend. Two preregistered studies then compare six contrastive encoders under a matched attack at three perturbation budgets. Robustly trained encoders (FARE, TeCoA, PMG, TRADES) transfer substantially more independent evidence than vanilla CLIP or SigLIP; all eight contrasts reject at the bootstrap floor. However, no cell reaches the blend-derived reference level. The three best cells fall within its replication band, leaving practical recovery undecided. Within robust encoders, per-sample alignment gain correlates with evidence gain ($ρ= 0.24-0.51$); within vanilla CLIP the correlation is consistent with zero. Across encoders we find no monotone alignment-evidence relation. VTS is therefore informative only within a fixed robust encoder, and we provide a reporting protocol in its place.
发表机构
- Singapore University of Technology and Design(新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。