发表机构
Sapienza University of Rome; University of Genoa; University of Trieste; University of Cagliari(罗马智慧大学; 热那亚大学; 的里雅斯特大学; 卡利亚里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型供应链安全问题,提出通过表示引导植入架构后门的攻击方法,该方法不影响训练数据等,通过触发控制模型表示转向攻击者目标,评估表明其损害模型多项性能,还提出审计防御方法。
AI 中文摘要
视觉语言模型(VLM)越来越多地通过模型供应链进行部署,第三方会分发预训练检查点、架构定义、文本编码器和导出的计算图,并在下游服务中复用。这种复用模型产生了一个对安全至关重要的信任边界。本文表明恶意提供者可通过表示引导在VLM供应链中植入架构后门。攻击通过对中间表示进行触发门控加法修改将休眠引导逻辑引入模型架构,不毒害训练数据、控制下游微调或在部署时修改提示。触发不存在时,修改为零,模型正常计算;触发存在时,引导方向将内部表示转向攻击者定义的目标。评估了多个VLM系列和下游任务,结果表明该后门损害了完整性、安全执行和排名公平性,同时在干净输入上保持正常行为。还表明共享VLM工件可携带针对下游服务的休眠引导逻辑,并提出一种审计防御方法,检查与模型工件一起分发的可执行逻辑而非仅其学习权重。
英文摘要
Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text encoders, and exported computation graphs are distributed by third parties and reused across downstream services. This reuse model creates a security-critical trust boundary: VLM deployments inherit not only learned parameters but also executable behavior encoded in shared model artifacts. In this paper, we show that a malicious provider can exploit this trust boundary by embedding architectural backdoors into VLM supply chains through representation steering. Our attack introduces dormant steering logic into the model architecture through a trigger-gated additive modification of an intermediate representation, without poisoning training data, controlling downstream fine-tuning, or modifying prompts at deployment time. When the trigger is absent, the modification reduces to zero and the model follows its normal computation, preserving clean utility. When the trigger is present, a steering direction shifts the internal representation toward an attacker-defined objective. We evaluate the attack across multiple VLM families and downstream tasks, including visual question answering, text-to-image generation, retrieval, and semantic response biasing. The results show that the proposed architectural steering backdoor compromises integrity, safety enforcement, and ranking fairness while preserving normal behavior on clean inputs. We further show that shared VLM artifacts can carry dormant steering logic against downstream services, and we propose an auditing defense that inspects the executable logic distributed with model artifacts rather than only their learned weights.