发表机构
College of Computing and Data Science, Nanyang Technological University; Center for Frontier AI Research, Agency for Science, Technology and Research (A*STAR)(南洋理工大学计算与数据科学学院; 新加坡科技研究局前沿人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言模型对抗鲁棒性不足问题,提出CARE框架,通过维护并协同微调多个鲁棒模型专家,实现知识融合,在多任务上性能优于单独训练的专家。
AI 中文摘要
CLIP等视觉-语言模型(VLMs)易受对抗攻击,给实际应用与部署带来严重问题。对抗微调是突出的防御方法,但不同微调策略常生成具有不同鲁棒性特征的专用模型,每种微调模型在部分评估场景表现优异,在其他场景则性能不佳,限制了防御能力。本文将这些专用微调模型称为鲁棒模型专家,提出协同对抗鲁棒性微调框架CARE(Collaborative Adversarial Robustness fine-tuning using Embedding alignment)。CARE在训练过程中维护多个专家模型,通过嵌入空间协调实现知识交换,并将所学知识整合为单个统一的鲁棒模型。专家模型相互受益,同时保留各自的专长,使最终模型继承互补的鲁棒性特性。本文在两种具有互补鲁棒性行为的对抗微调策略上验证了CARE,在经典图像分类和下游视觉-语言任务上的大量实验显示了该方法的有效性,CARE的性能优于单独训练的模型专家,结果表明跨模型专家的协同学习是提升对抗鲁棒性的有前景方向。
英文摘要
Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment. Adversarial fine-tuning emerges as a prominent defense method; however, different fine-tuning strategies often produce specialized models with distinct robustness characteristics. Each fine-tuned model in turn thrives in some evaluation settings but falters on others, limiting their defensive capabilities. We refer to these specialized fine-tuned models as robust model experts and propose a collaborative adversarial fine-tuning framework: CARE - Collaborative Adversarial Robustness fine-tuning using Embedding alignment. CARE maintains multiple experts during training, enables knowledge exchange through embedding-space harmonization, and consolidates the learned knowledge into a single unified robust model. Experts benefit from one another while preserving their individual specializations, enabling the final model to inherit complementary robustness properties. In this paper, we demonstrate CARE on two different adversarial fine-tuning strategies with complementary robustness behaviors. Extensive experiments on classic image classification and downstream vision-language tasks display the effectiveness of our approach, with CARE being able to outperform individually learned model experts. The results suggest that collaborative learning across model experts is a promising direction for improving adversarial robustness.