AI 中文总结
该研究提出一种向VLM植入可编程后门的方法,通过启发式投毒策略和特征空间触发器隐写,可在推理时动态选择未见过的目标字幕并触发对应输出,攻击成功率高且能规避部分防御。
AI 中文摘要
现有视觉语言模型(VLM)的后门通常被视为静态漏洞:一对一和N对N攻击在受害者模型训练前将一个或多个触发器绑定到有限的目标集合,这一假设严重低估了威胁。本文展示,单次投毒阶段即可向VLM植入可编程后门,使攻击者在推理时可选择此前未见过的目标字幕语义,并按需合成对应的隐蔽触发器。与固定映射攻击不同,所提的任意对任意字幕控制范式将训练后目标选择与投毒解耦,无需重新训练VLM即可动态控制目标字幕。该方法包含两个组件:其一,启发式投毒策略让模型接触多样化的触发器-字幕对,促使其学习通用的“触发器作为指令”规则,而非记忆特定后门模式;其二,特征空间触发器隐写方法将攻击者指定的任意目标字幕映射为隐蔽视觉触发器,可实现为范数控制扰动或非语义补丁。将这些触发器插入任意图像后,中毒的VLM会生成与所选目标字幕语义一致的输出,即便该目标在投毒阶段未出现。大量实验表明,本文攻击实现了高任意对任意字幕控制成功率,保留了干净模型的效用,且在多种经典后门防御下仍有效。
英文摘要
Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers to a finite set of targets before victim training. This assumption substantially underestimates the threat. We show that a single poisoning phase can implant a programmable backdoor into a VLM, allowing an attacker to choose previously unseen target-caption semantics at inference time and synthesize corresponding stealthy triggers on demand. Unlike fixed-mapping attacks, the proposed any-to-any caption-control paradigm decouples post-training target selection from poisoning, enabling dynamic control of target captions without retraining the VLM. Our method has two components. First, a heuristic poisoning strategy exposes the model to diverse trigger-caption pairs, encouraging it to learn a general trigger-as-instruction rule rather than memorize a specific backdoor pattern. Second, a feature-space trigger steganography method maps any attacker-specified target caption to a stealthy visual trigger, implemented as either a norm-controlled perturbation or a non-semantic patch. Once inserted into arbitrary images, these triggers cause the poisoned VLM to generate outputs semantically aligned with the chosen target caption, even when the target was unseen during poisoning. Extensive experiments show that our attack achieves high any-to-any caption-control success rates, preserves clean model utility, and remains effective under several classical backdoor defenses.