发表机构
Chongqing University; Huazhong University of Science and Technology; City University of Hong Kong; Griffith University(重庆大学; 华中科技大学; 香港城市大学; 格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 AdvPCS,一种针对可提示概念分割的通用跨提示对抗攻击,通过最小-最大提示优化、全局-局部感知欺骗和时间记忆错位攻击,生成单一 UAP 在多种提示下显著降低模型性能。
AI 中文摘要
Segment Anything Model (SAM) 在视觉分割领域取得了显著性能。最新的 SAM3 将可提示分割扩展到概念级预测,拓宽了分割基础模型的应用范围。尽管近期研究揭示了 SAM 和 SAM2 易受对抗样本攻击,但 SAM3 在概念分割范式下的鲁棒性尚未被探索。此外,现有针对 SAM 系列模型的对抗攻击在跨提示迁移性方面表现有限。为此,我们提出了 AdvPCS,一种针对可提示概念分割 (PCS) 的通用跨提示对抗攻击,包括最小-最大提示优化策略、全局-局部感知欺骗攻击和时间转换偏差攻击。具体而言,我们首先通过最小-最大双层优化识别最难攻击的提示。在内层最大化中,我们增强候选点、框和文本提示的多样性。在外层最小化中,我们根据检测器输出的置信度分数选择响应最高的提示。给定所选提示,我们应用感知欺骗攻击以在联合提示下最小化全局和局部存在概率,并采用时间记忆错位攻击以最大化帧间语义不一致并破坏记忆指针。在四个基准数据集上的大量实验表明,我们方法生成的单一通用对抗扰动 (UAP) 能够跨不同视频的帧进行泛化,并在点、框和文本提示下实现强攻击性能。特别地,在文本提示下,它将 SA-CO 数据集上各种 PCS 模型的平均 mIoU 降至 5% 以下,展示了强大的攻击能力。
英文摘要
The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial examples, the robustness of SAM3 under the concept segmentation paradigm remains unexplored. In addition, existing adversarial attacks on SAM-series models exhibit limited cross-prompt transferability. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack for Promptable Concept Segmentation (PCS), including a min-max prompt optimization strategy, a global-local perception deception attack, and a temporal transition deviation attack. Specifically, we first identify the hardest-to-attack prompts via min-max bilevel optimization. In the inner maximization, we enhance diversity over candidate point, box, and text prompts. In the outer minimization, we select prompts with the highest responses based on the confidence scores output by the detector. Given the selected prompts, we apply the perception deception attack to minimize both global and local existence probabilities under joint prompting and employ the temporal memory misalignment attack to maximize inter-frame semantic inconsistency and corrupt memory pointers. Extensive experiments on four benchmark datasets show that a single universal adversarial perturbation (UAP) generated by our method generalizes across frames from different videos and achieves strong attack performance under point, box, and text prompts. In particular, it reduces the average mIoU of various PCS models on the SA-CO dataset to below 5% under text prompts, demonstrating strong attack ability.
CommentsAccepted by NeurIPS 2026