Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models
对抗攻击已揭示答案:面向视觉语言模型的定向偏差引导测试时防御
机构 * University of Science and Technology of China(中国科学技术大学) ; National Key Laboratory of Deep Space Exploration, Deep Space Exploration Laboratory(国家深空探测重点实验室,深空探测实验室) ; The Chinese University of Hong Kong(香港中文大学) ; Zhejiang University(浙江大学) ; Ant Group(蚂蚁集团)
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV
AI总结 提出定向偏差引导防御(DBD),利用对抗样本在CLIP特征空间中沿主导方向偏移的现象,通过估计防御方向并采用DB分数双流重建策略恢复鲁棒表示,在15个数据集上实现最先进对抗鲁棒性且保持干净准确率。
Comments Accepted by ICLR2026