LoRango:两个LoRA解锁扩散模型中的隐藏行为
LoRango: It Takes Two LoRAs to Unlock Hidden Behaviors in Diffusion Models
- Fudan University(复旦大学)
- Xi’an Jiaotong-Liverpool University(西交利物浦大学)
- Purple Mountain Laboratories(紫金山实验室)
- Worcester Polytechnic Institute(伍斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出LoRango攻击,利用成对LoRA适配器在扩散模型中触发隐藏行为,实现高成功率,揭示单独检查适配器的安全不足。
AI中文摘要:
用户通常组合多个低秩适应(LoRA)适配器,以使用不同的主题、风格和视觉属性来个性化图像。然而,单独检查适配器并不能确定其组合的安全性。我们识别并描述了一种文本到图像扩散中的成对条件攻击:单独有用且外观良性的适配器在与特定匹配的伙伴(其身份作为触发器)共同加载时,会重定向图像生成。我们引入LoRango,通过互补的签名(Signature)和载荷(Payload)适配器实现此攻击。签名适配器将成对特定代码写入中间载体表示,而载荷适配器使用代码选择性响应和相反的信号/参考分支。这些分支在单独适配器和不匹配对中近似抵消;匹配的代码读取器对齐在原生GEGLU块内打破抵消并释放编程动作。两个适配器均导出为与标准加载器兼容的普通静态LoRA文件,无需提示触发器或基础流水线修改。LoRango在SD v1.5上实现97.9%的匹配对攻击成功率,在SDXL上实现98.7%,而单独加载植入适配器时仅为2.8-4.6%。进一步实验评估了配对选择性、单独保真度、对部署变化的鲁棒性以及跨去噪器架构的适用性。这些发现表明,单独适配器检查不足以评估多LoRA个性化的安全性,并激励对适配器组合进行审计。
英文摘要:
Users commonly combine multiple Low-Rank Adaptation (LoRA) adapters to personalize images with different subjects, styles, and visual attributes. Yet inspecting adapters individually does not establish the safety of their composition. We identify and characterize a pair-conditioned attack in text-to-image diffusion: individually useful and benign-appearing adapters redirect image generation when co-loaded with a specifically matched partner, whose identity serves as the trigger. We introduce LoRango to realize this attack through complementary Signature and Payload adapters. The Signature writes a pair-specific code into intermediate carrier representations, while the Payload uses code-selective responses and opposing signal/reference branches. These branches approximately cancel for standalone adapters and mismatched pairs; matched code-reader alignment breaks cancellation within native GEGLU blocks and releases the programmed action. Both adapters are exported as ordinary static LoRA files compatible with standard loaders, requiring no prompt trigger or base-pipeline modification. LoRango achieves matched-pair attack success rates of 97.9\% on SD v1.5 and 98.7\% on SDXL, compared with 2.8--4.6\% when implanted adapters are loaded individually. Further experiments evaluate pair selectivity, standalone fidelity, robustness to deployment variations, and applicability across denoiser architectures. These findings show that individual-adapter inspection is insufficient to assess the security of multi-LoRA personalization and motivate auditing adapter compositions.