Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
高质量文本,稳健视觉:语言在增强视觉语言模型视觉稳健性中的作用
机构 * The University of Tokyo(东京大学) ; National Institute of Informatics(日本信息处理研究所) ; The University of Tokyo, National Institute of Informatics(东京大学、日本信息处理研究所)
AI总结 研究针对视觉语言模型对抗攻击问题,提出质量文本引导的对抗微调方法QT-AFT,利用高质量字幕提升视觉编码器对抗鲁棒性,在多个零样本数据集上实现了最先进的零样本对抗鲁棒性和纯净准确率,并揭示了语言增强视觉鲁棒性的关键见解。
Comments ACMMM 2025 Accepted. Codes are available at: https://github.com/futakw/QTAFT