发表机构
The University of Queensland; Zhejiang University; Southeast University(昆士兰大学; 浙江大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出OrthoPurify,通过一步正交投影移除被劫持的方向,净化后门LVLM权重,无需重训练或推理开销,将攻击成功率降至接近零并保持性能。
AI 中文摘要
大型视觉-语言模型(LVLMs)正越来越多地部署在安全关键应用中,然而它们仍然容易受到后门攻击。防御此类攻击仍然代价高昂,因为现有方法要么需要在干净数据上进行大量重训练,要么需要在推理时进行逐查询干预。为了解决这一局限,我们提出了OrthoPurify,一种通过一步正交投影来净化被植入后门模型权重的更高效方法。具体而言,通过对后门权重更新的结构分析,我们发现后门是通过将少量权重更新方向从任务适应转向后门捷径编码来编码的,我们将这一现象称为方向劫持。然而,识别这些被劫持的方向需要一个良性参考模型,而防御者通常无法获得该模型。我们证明,通过在仅少量干净样本上微调预训练权重而获得的伪良性模型,提供了足够的近似,因为主导更新方向在前几步梯度内趋于稳定。OrthoPurify利用这一伪良性参考来隔离被劫持的方向,并通过在权重更新上进行一次投影来移除它们。大量实验表明,OrthoPurify将攻击成功率降至接近零,同时在各种基准上保持原有性能,无需重训练被植入后门的模型,也不引入推理时的开销。我们的代码可在以下网址公开获取:此https URL。
英文摘要
Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks. Defending against such attacks remains costly, as existing methods require either extensive retraining on clean data or per-query intervention at inference time. To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection. Specifically, through structural analysis of backdoor weight updates, we find that the backdoor is encoded by diverting a small number of weight update directions from task adaptation to backdoor shortcut encoding, a phenomenon we term direction hijacking. However, identifying these hijacked directions requires a benign reference model, which is typically inaccessible to the defender. We show that a pseudo-benign model, obtained by fine-tuning the pretrained weights on only a small set of clean samples, provides a sufficient approximation, as the dominant update directions stabilize within the first few gradient steps. OrthoPurify uses this pseudo-benign reference to isolate the hijacked directions and removes them through a single projection on the weight update. Extensive experiments show that OrthoPurify reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead. Our code is publicly available at https://github.com/womeimingzi/OrthoPurify.
Comments25 pages, 9 figures, 14 tables