arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高斯核心LoRA:面向广泛概念擦除的分布感知动态适配

Gaussian Core LoRA: Distribution-Aware Dynamic Adaptation for Broad Concept Erasure

Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang, Hua Meng, Zhengchun Zhou

arXiv 2609.01433首次发表:更新:

发表机构

Southwest Jiaotong University; University of Electronic Science and Technology of China(西南交通大学; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有LoRA等概念擦除方法对复杂目标概念擦除不足或过度编辑的问题,提出高斯核心LoRA框架,通过高斯混合模型实现原型自适应擦除,在多指标上优于基线且兼容主流模型。

AI 中文摘要

概念擦旨在抑制文本到图像扩散模型中不安全、隐私敏感或不合需要的生成内容,同时保留良性语义、视觉质量和部署效率。现有的基于适配器的方法,如低秩适配(Low-Rank Adaptation,LoRA),通常冻结扩散主干并学习轻量级参数更新,以引导生成远离目标语义。然而,这些方法通常为每个目标概念分配静态语义擦除方向,该假设对于广泛且复杂的目标概念过于粗糙,因为一个概念往往包含多个潜在语义原型,涉及不同对象、场景或关系,需要不同的局部擦除方向。单一LoRA更新会平均这些异质擦除需求,导致对困难原型的擦除不足以及对附近良性语义的过度编辑。为解决此限制,我们提出高斯核心LoRA(Gaussian Core LoRA),一种分布感知的低秩适配框架。它在提示特征空间中拟合高斯混合模型,以估计目标概念内的潜在语义原型。推理期间,每个输入提示被投影到该特征空间以计算其高斯后验责任,这些责任条件化核心生成器,生成共享LoRA秩空间的提示特定、范数有界的残差重构。这使得使用单一轻量级适配器实现原型自适应擦除成为可能。与各指标上最强基线相比,高斯核心LoRA将平均攻击成功率(Attack Success Rate,ASR)降低7.95%,将COCO Fréchet Inception距离(FID)降低14.72%,并将CLIP分数提高4.98%。进一步实验表明,该方法对对抗性提示具有鲁棒性,可扩展至多身份和多风格擦除,且与SDXL和FLUX兼容。

英文摘要

Concept erasure aims to suppress unsafe, privacy-sensitive, or undesirable generations in text-to-image diffusion models while preserving benign semantics, visual quality, and deployment efficiency. Existing adapter-based methods, such as Low-Rank Adaptation (LoRA), typically freeze the diffusion backbone and learn lightweight parameter updates to steer generation away from target semantics. However, these methods usually assign a static semantic erasure direction to each target concept. This assumption is overly coarse for broad and complex target concepts, since a concept often contains multiple latent semantic prototypes involving different objects, scenes, or relations, and requires different local erasure directions. A single LoRA update averages these heterogeneous erasure demands, leading to under-erasure on difficult prototypes and over-editing of nearby benign semantics. To address this limitation, we propose Gaussian Core LoRA, a distribution-aware low-rank adaptation framework. It fits a Gaussian mixture model in the prompt feature space to estimate latent semantic prototypes within the target concept. During inference, each input prompt is projected into this feature space to compute its Gaussian posterior responsibilities, which condition the core generator to produce a prompt-specific, norm-bounded residual reconfiguration of the shared LoRA rank space. This enables prototype-adaptive erasure with a single lightweight adapter. Compared with the strongest baseline on each metric, Gaussian Core LoRA reduces average Attack Success Rate (ASR) by 7.95%, lowers COCO Fr'echet Inception Distance (FID) by 14.72%, and improves CLIP Score by 4.98%. Further experiments show robustness to adversarial prompts, scalability to multi-identity and multi-style erasure, and compatibility with SDXL and FLUX.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑