发表机构
College of Computing and Data Science, Nanyang Technological University; Center for Frontier AI Research, Agency for Science, Technology and Research (A*STAR); School of Electrical Engineering and Automation, Fuzhou University; ByteDance(南洋理工大学计算与数据科学学院; 新加坡科技研究局前沿人工智能研究中心; 福州大学电气工程与自动化学院; 字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言模型易受对抗扰动的问题,提出协同演化跨模态攻击框架,联合优化文本与视觉空间,在多任务上展现出强攻击性能与跨任务可迁移性。
AI 中文摘要
视觉-语言模型(VLMs)在多模态任务中展现出强大的泛化能力,但仍易受对抗扰动影响。现有攻击通常遵循单轨迹梯度优化或特定任务目标,限制了搜索空间探索与跨任务可迁移性。我们提出一种演化计算引导的统一VLMs跨模态攻击框架,该框架自适应地在文本与视觉空间中搜索。在文本侧,它围绕源类别表示演化难负语义嵌入,以提供多样的跨模态排斥;在视觉侧,它维护一组目标区域扰动,并结合基于动量的梯度更新与演化选择、变异和交叉操作,以更可靠地探索多条可行轨迹。联合优化语义负引导与局部扰动,可生成对抗样本,使其在跨视觉-语言任务中持续将源对象语义转向目标类别。理论分析表明,与单轨迹优化相比,协同演化搜索保持了扰动可行性,防止了最优观测适应度下降,并提高了到达高间隔对抗区域的概率。在Florence-2、OFA和UnifiedIO-2上开展的实验,验证了其在图像字幕、目标检测、区域分类和目标定位等任务中均表现出强大的整体攻击性能。消融研究进一步证实了文本侧语义演化与图像侧扰动演化的互补有效性,以及该框架的效率与跨任务可迁移性。
英文摘要
Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attacks typically follow single-trajectory gradient optimization or task-specific objectives, limiting search-space exploration and cross-task transferability. We propose an evolutionary-computation-guided cross-modal attack framework for unified VLMs. The framework adaptively searches both textual and visual spaces. On the textual side, it evolves hard negative semantic embeddings around the source-category representation to provide diverse cross-modal repulsion. On the visual side, it maintains a population of object-region perturbations and combines momentum-based gradient updates with evolutionary selection, mutation, and crossover to more reliably explore multiple feasible trajectories. Jointly optimizing semantic negative guidance and localized perturbations generates adversarial examples that consistently shift source-object semantics toward target categories across vision-language tasks. Theoretical analyses show that the co-evolutionary search preserves perturbation feasibility, prevents degradation of the best observed fitness, and increases the probability of reaching high-margin adversarial regions compared with single-trajectory optimization. Experiments on Florence-2, OFA, and UnifiedIO-2 demonstrate strong overall attack performance across image captioning, object detection, region categorization, and object localization. Ablation studies further verify the complementary effectiveness of text-side semantic evolution and image-side perturbation evolution, as well as the framework's efficiency and cross-task transferability.
Comments15 pages, 7 figures, and 8 tables; includes supplementary material