arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31878cs.CR

TokenScanner:通过全词汇扫描检测文本到图像LoRA中的后门并发现触发器

TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning

Boliang Liu, Jing Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

TokenScanner通过扫描全词汇并测量LoRA响应,检测文本到图像扩散模型中的后门LoRA并发现触发器,在840个后门和840个良性LoRA上达到95.96% AUC和99.40% Hit@5。

中文摘要 AI 辅助

LoRA被广泛研究用于适配基础文本到图像扩散模型。然而,一个被植入后门的LoRA可以隐藏后门:它在大多数情况下表现正常,但当输入提示中出现隐藏的后门触发器时,会产生攻击者指定的内容(后门目标)。在使用不受信任的LoRA之前检测此类后门对于LoRA适配的安全性至关重要。我们提出了TokenScanner,一种针对LoRA微调文本到图像扩散模型的模型级词汇扫描器,旨在发现恶意触发器以实现可信的LoRA适配。关键观察是,在训练提示中占较大比例出现的后门触发器标记往往比无关标记引发更显著的标记特定LoRA响应。因此,TokenScanner扫描分词器词汇表,并测量U-Net和文本编码器中逐标记的LoRA响应。它利用这些响应来检测被植入后门的LoRA,并对候选触发器标记进行排序,以便后续测试后门激活。在七个后门设置上的实验(包括840个被植入后门的LoRA和840个真实世界良性测试LoRA)表明,TokenScanner在FPR为10.95%时实现了95.96%的AUC和96.55%的TPR。在所有七个设置中,触发器发现还实现了89.40%的Hit@1和99.40%的Hit@5。

英文摘要

LoRAs are widely studied for adapting base text-to-image diffusion models. However, a backdoored LoRA can hide a backdoor: it behaves normally in most cases, but produces attacker-specified content (the backdoor target) when a hidden backdoor trigger appears in the input prompt. Detecting such backdoors before using an untrusted LoRA is important for the safety of LoRA adaptation. We present TokenScanner, a model-level vocabulary scanner for backdoor detection within a LoRA fine-tuned text-to-image diffusion model, aiming to discover the malicious trigger for trustworthy LoRA adaptation. The key observation is that backdoor trigger tokens that appear in a larger proportion of training prompts tend to induce more prominent token-specific LoRA responses than those induced by unrelated tokens. TokenScanner therefore scans the tokenizer vocabulary and measures token-wise LoRA responses in the U-Net and the text encoder. It uses these responses to detect backdoored LoRAs and rank candidate trigger tokens for subsequent testing of backdoor activation. Experiments on seven backdoor settings, comprising 840 backdoored LoRAs and 840 real-world benign test LoRAs, show that TokenScanner achieves 95.96% AUC and 96.55% TPR at an FPR of 10.95%. It also achieves 89.40% Hit@1 and 99.40% Hit@5 for trigger discovery across all seven settings.

发表机构

  • The Australian National University(澳大利亚国立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑