AI 中文总结
ReACT-CLIP是无需训练的响应感知测试时防御,通过跨噪声特征漂移与预测不稳定性得分动态调整修正强度,在多数据集上提升了CLIP类模型的对抗鲁棒性且保留干净准确率。
AI 中文摘要
无需训练的测试时防御是提升CLIP类视觉-语言模型对抗鲁棒性的实用方法,且无需修改预训练模型。但现有防御的修正强度通常针对窄范围攻击预算固定,而推理时攻击预算未知,不同样本所需修正程度不同,这种不匹配会导致防御随攻击增强而大幅失效。本文提出ReACT-CLIP,一种响应条件测试时防御,可分别确定每个输入的修正强度及是否需要防御干预。核心发现:低噪声与高噪声探测间CLIP视觉特征漂移的相对增量,是样本特定的修正需求代理指标。ReACT-CLIP将这种跨噪声相对漂移映射到高斯噪声尺度,用于构建稳定的噪声平均特征锚点,使修正范围适配每个输入。为判断是否需要干预,进一步发现:干净输入在弱空间增广下保持稳定的类别概率分布,而对抗输入则表现出更大变化,ReACT-CLIP用Jensen-Shannon散度计算预测不稳定性得分来量化这种变化,并将其与跨噪声相对漂移结合形成防御干预得分。ReACT-CLIP无需模型或提示训练,其修正强度映射仅校准一次,可跨数据集和攻击预算使用。在12个下游数据集、ImageNet及其分布偏移变体上,ReACT-CLIP在各类攻击类型和强度下均实现了显著的鲁棒性提升,同时基本保持干净准确率。
英文摘要
Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However, their correction strength is typically fixed for a narrow range of attack budgets, even though the attack budget is unknown at inference and the required correction varies across samples. We show that this mismatch causes existing defenses to degrade sharply as attacks strengthen. We introduce ReACT-CLIP, a response-conditioned test-time defense that separately determines how strongly each input should be corrected and whether defensive intervention is necessary. Our key observation is that the relative increase in CLIP visual-feature drift between low- and high-noise probes provides a graded, sample-specific proxy for correction demand. ReACT-CLIP maps this relative cross-noise drift to the Gaussian noise scale used to construct a stable, noise-averaged feature anchor, enabling the corrective reach to adapt to each input. To determine whether intervention is necessary, we further observe that clean inputs retain stable class-probability distributions under weak spatial augmentations, whereas adversarial inputs exhibit greater variation. ReACT-CLIP quantifies this variation using a prediction-instability score computed by Jensen--Shannon divergence and combines it with relative cross-noise drift to form the defensive intervention score. ReACT-CLIP requires no model or prompt training, and its correction-strength mapping is calibrated once and fixed across datasets and attack budgets. Across 12 downstream datasets, as well as ImageNet and its distribution-shifted variants, ReACT-CLIP delivers substantial robustness gains across diverse attack types and strengths while largely preserving clean accuracy.