arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06795cs.CRcs.AIcs.CL

LoRAScan:通过下投影激活尖峰检测大语言模型低秩适配器中的后门提示

LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes

  • Michigan Technological University(密歇根理工大学)
  • Miami University(迈阿密大学)
  • Lehigh University(理海大学)

机构由 AI 辅助整理,请以论文原文为准。

Doniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen Yang

AI总结:

LoRAScan是首个无需修改适配器参数、在推理时检测并拒绝含触发词输入的适配器感知防御方法,在LLM后门基准测试中拒98.49%恶意输入,性能优于现有方法。

AI中文摘要:

低秩适配(LoRA)通过紧凑适配器实现大语言模型的高效专业化与分发,但不可信适配器会引入供应链威胁:植入后门的适配器会在输入包含隐藏触发词时,使模型生成有害内容、恶意代码、政治宣传或隐蔽广告。与适配器无关的防御方法会将适配器与基础模型合并,这会稀释后门信号并降低检测性能。现有的适配器感知方法要么训练防御性适配器修复被后门植入的基础模型(解决反向问题而非保护适配器本身),要么依赖分类器将整个适配器标记为可疑并需单独缓解,这些方法均未考虑后门适配器中含触发词的输入所产生的独特潜在空间特征。我们提出LoRAScan,这是首个在推理时检测并拒绝含触发词输入的适配器感知防御方法,无需修改适配器参数。我们的关键发现是,约5%的LoRA插入位点在干净输入下保持稳定,但在触发词存在时,LoRA下投影激活会出现高度集中的尖峰;LoRAScan在模型部署前识别这些低方差插入位点,并在推理时对其进行监测。在标准LLM后门基准测试中,LoRAScan拒绝约98.49%的恶意输入,且对干净输入的错误率较低,在多种评估场景下均优于现有防御方法。

英文摘要:

Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, untrusted adapters introduce a supply-chain threat: a backdoored adapter can cause a model to generate harmful content, malicious code, political propaganda, or covert advertisements when an input contains a hidden trigger. Adapter-agnostic defenses merge the adapter with the base model, which dilutes backdoor signals and reduces detection performance. Existing adapter-aware methods do not address how to safely use a potentially backdoored adapter. Instead, they either train a defensive adapter to repair a backdoored base model, addressing the inverse problem rather than securing the adapter itself, or rely on a classifier that flags the entire adapter as suspicious and requires separate mitigation. These methods overlook the distinct latent-space signatures produced by trigger-bearing inputs in backdoored adapters. We introduce LoRAScan, the first adapter-aware defense that detects and rejects trigger-bearing inputs at inference time without modifying adapter parameters. Our key observation is that a small subset of LoRA insertion sites, approximately 5%, remains stable across clean inputs but exhibits highly concentrated spikes in LoRA down-projection activations when a trigger is present. LoRAScan identifies these low-variance insertion sites before model deployment and monitors them during inference. Across standard LLM backdoor benchmarks, LoRAScan rejects approximately 98.49 of malicious inputs with a small error rate on clean inputs, outperforming existing defenses across diverse evaluation settings.

↑