发表机构
Michigan State University; JPMorgan AI Research; Henry Ford Health(密歇根州立大学; 摩根大通人工智能研究院; 亨利福特医疗集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对答案保留型攻击,发现视觉-语言系统的置信度信号不具备固有鲁棒性,四类防御均无效,协调攻击可大幅降低置信度门控下的接受准确率。
AI 中文摘要
部署的视觉-语言系统常基于置信度对输出答案进行门控,这使得置信度鲁棒性与监管工作密切相关。我们研究仅保留生成答案字节完全一致的白盒、仅图像攻击下的置信度读数。在可达性假设下,不可移动的读数无法优于此前的答案字符串准确率,其汇总值为0.617。独立于该假设,低于可测量阈值的均匀幅度证书保证对抗性判别能力不低于同一基准。在四个视觉-语言模型、三个视觉问答基准、五个已部署置信度通道和两个防御估计器中,直接或针对代理的攻击产生的逐元素可行扰动,在全部84种估计器-单元组合中均否定了该均匀证书。协调的、感知正确标签的攻击,在全部60种已部署通道单元中,将对抗性判别能力降至答案字符串基准或其以下,包括全部59种初始高于该基准的单元。隐藏状态干预和开放式文本模型激活空间复制表明,可在表示层面而非仅通过对抗图像诱导可比的置信度变化。测试的四类防御均未在其特定评估下建立鲁棒替代方案。在置信度门控模拟中,针对令牌概率的协调攻击转移至隐藏状态门,导致多达84.8%此前被拒绝的错误答案被接受。按各基准的自然正确率 prevalence 重新加权后,在转移场景下12个单元中有8个、直接针对门的攻击下全部12个单元中,接受准确率低于无门基线。因此,在研究的威胁模型和预算下,置信度是对完整性敏感而非固有鲁棒的监管信号。
英文摘要
Vision-language models are increasingly deployed behind a confidence gate: the system reads how confident the model is in its answer and defers when confidence is low. This makes the confidence signal itself worth attacking. We show that a white-box adversary who perturbs only the input image, within an L-infinity budget of 8/255 and while keeping the model's answer byte-identical, can invert the confidence ranking, lowering it on correct answers and raising it on wrong ones until the signal points the wrong way. Most of the inversion persists even when the whole next-token distribution is held near the clean one, so the answer does not determine the confidence attached to it. Confidence is a separate signal read from the same network, and it can be corrupted on its own. A gate reading it is turned against itself, rejecting good answers and accepting wrong ones it was built to catch. Across four vision-language models and three visual question-answering benchmarks, the attack drives the model-internal readouts below chance in 83 of 84 readout-by-cell profiles under an adversary that knows which answers are correct; for the two readouts carrying a disjoint calibration reference, it falls below chance under an adversary that does not. Training a probe on frozen hidden states does not fix this: the robustness it gains is paid for with the information that made it useful. Nor does reading confidence from a separate, independently trained model, which holds up only until the attacker reaches it and then falls into the same regime. How far an answer-preserving adversary can reach a signal governs where it survives; whether a robust and informative readout can be built remains open. For deployment, a gate under this attack can admit most wrong answers it would otherwise catch and, corrected for how often the model is wrong, can leave the system worse off than using no gate at all.