arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CrACK:针对协作视觉基础模型中跨模型一致性的对抗攻击

CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models

Feifei Liu, Jintao Cheng, Chi Man Vong, Xiaoyu Tang

arXiv 2609.07499首次发表:更新:

发表机构

South China Normal University; Hong Kong University of Science and Technology; University of Macau(华南师范大学; 香港科技大学; 澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CrACK攻击,利用协作视觉基础模型间未验证语义一致性的接口,通过两阶段方法在推理时破坏跨模态亲和矩阵和操纵预测,导致系统性能灾难性退化,证明模型间接口必须视为安全边界。

AI 中文摘要

无需训练的协作流水线集成了CLIP、SAM和DINO等视觉基础模型,实现了强大的开放词汇密集预测,并越来越多地部署在安全关键应用中。这些系统的安全性通常被认为源于其单个模型的鲁棒性。我们挑战了这一假设。我们识别出所有协作流水线共有的一个漏洞:每个模型都消费另一个模型的中间输出,而不验证语义一致性,我们将这一未经证实的假设称为语义-空间对齐依赖。现有的对抗攻击针对单个模型,忽视了这一前提,导致模型间接口完全不受防护。我们提出了CrACK(跨模型对抗一致性攻击),这是一种推理时攻击,利用该接口而无需修改任何输入像素、模型权重或训练数据。CrACK分两个阶段运作:对抗亲和性矛盾注入通过CLIP补丁级语义引导下反转SAM编码器特征来破坏跨模态亲和矩阵,语义接口投毒则通过从CLIP文本嵌入导出的最大距离标签排列来操纵预测。在四个协作流水线和八个基准上的实验表明,CrACK导致灾难性退化,而每个单独模型继续产生其不变的独立输出,使基于单模型的防御在结构上失明。这种破坏进一步级联到大视觉语言模型推理中,促使LLaVA等模型从视觉上完整的输入产生错误响应。我们的结果表明,协作AI系统的安全性不能简化为其组件的鲁棒性,模型间特征接口必须被视为一等安全边界。

英文摘要

Training-free collaborative pipelines that integrate Vision Foundation Models such as CLIP, SAM, and DINO achieve strong open-vocabulary dense prediction and are increasingly deployed in safety-critical applications. The security of these systems is commonly assumed to follow from the robustness of their individual models. We challenge this assumption. We identify a vulnerability shared by every collaborative pipeline: each model consumes the intermediate output of another without verifying semantic consistency, an unverified premise that we term the semantic-spatial alignment dependency. Existing adversarial attacks target a single model and overlook this premise, leaving the inter-model interface entirely unguarded. We propose CrACK (Cross-model Adversarial Consistency attack), an inference-time attack that exploits this interface without modifying any input pixel, model weight, or training data. CrACK operates in two stages: Adversarial Affinity Contradiction Injection corrupts the cross-modal affinity matrix by inverting SAM encoder features under the guidance of CLIP patch-level semantics, and Semantic Interface Poisoning steers the prediction through a max-distance label permutation derived from CLIP text embeddings. Experiments on four collaborative pipelines across eight benchmarks show that CrACK causes catastrophic degradation while every individual model continues to produce its unchanged standalone output, rendering per-model defenses structurally blind. The corruption further cascades into large vision-language model reasoning, driving models such as LLaVA to produce erroneous responses from visually intact inputs. Our results show that the security of a collaborative AI system cannot be reduced to the robustness of its components, and that inter-model feature interfaces must be treated as first-class security boundaries.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑