发表机构
Southeast University; The University of Queensland; University of Sydney; A*STAR(东南大学; 昆士兰大学; 悉尼大学; A*STAR)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PSR防御方法,利用投影脆弱性识别并恢复带后门视觉-语言模型中的敏感通道,将攻击成功率降至近零且不增加推理开销。
AI 中文摘要
视觉-语言模型(VLMs)展现出强大的多模态能力,但仍易受到通过中毒微调数据植入的后门攻击。现有防御方法通常需要在微调期间进行大量参数更新,或在推理时产生每次查询的额外开销。为解决这些局限性,我们提出了扰动-选择-恢复(PSR),一种训练后防御方法,对投影接口进行稀疏更新,且在推理时不引入额外计算。我们揭示,带后门的VLM投影器对有限扰动的敏感性显著高于干净的VLM投影器,我们将这一现象称为投影脆弱性。基于此发现,PSR识别带后门VLM每个投影层中对扰动最敏感的输出通道,并将其参数恢复为相应的预训练值。跨多个任务的实验表明,PSR将攻击成功率降至接近零,同时保持干净任务的性能。
英文摘要
Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM projectors, a phenomenon we term projection fragility. Building on this finding, PSR identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values. Experiments across multiple tasks show that PSR reduces attack success rates to near zero while preserving clean-task performance.
Comments14 pages, 4 figures