arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过测试时提示自适应增强视觉语言模型的鲁棒性

Robustifying Vision-Language Models via Test-Time Prompt Adaptation

Xingyu Zhu, Huanshen Wu, Shuo Wang, Beier Zhu, Jiannan Ge, Jiaheng Zhang, Long Chen

arXiv 2607.09450首次发表:更新:

发表机构

University of Science; National University of Singapore; The Hong Kong University of Science(科学大学; 新加坡国立大学; 香港科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对对抗扰动下视觉语言模型性能下降问题,提出RITA框架,通过最优传输实现分布级对齐,引入动态缓存积累线索,提升了模型对抗鲁棒性且不损干净准确率。

AI 中文摘要

预训练的视觉语言模型(如CLIP)在零样本泛化方面表现出色,但在对抗性扰动下性能会急剧下降。现有测试时自适应方法通常依赖样本级置信启发式,忽略了数据的内在分布结构。本文提出RITA框架,从样本级估计转向分布级对齐。具体而言,利用最优传输使增强视觉特征分布与文本原型对齐,减轻对抗性异常值并纠正跨模态语义错位。还引入动态缓存逐步积累可靠线索进行在线优化。实验表明RITA显著提高对抗鲁棒性且不影响干净准确率。

英文摘要

Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional structure of the data. This sample-centric approach limits robustness, as it fails to distinguish confident adversarial mispredictions from true semantic consistency. In this work, we observe that adversarial distortion is structurally brittle: while holistic representations are corrupted, semantic integrity is often preserved in the distribution of augmented views. Motivated by this insight, we propose RITA, a Robust test-tIme prompt-TAdaptation framework that shifts from sample-level estimates to distribution-level alignment. Specifically, RITA employs optimal transport to align the distribution of augmented visual features with textual prototypes, mitigating adversarial outliers and rectifying cross-modal semantic misalignment. Furthermore, we introduce a dynamic cache to progressively accumulate reliable cues from the test stream for online refinement. Extensive experiments demonstrate that RITA significantly improves adversarial robustness without compromising clean accuracy.

CommentsICML 2026 regular

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑