arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiffSafeMerge:缓解扩散模型合并中的后门继承问题

DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging

Jiayang Zhang, Ji Guo, Jiachen Li, Wenshu Fan, Wenbo Jiang

arXiv 2608.09445首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Wuhan University of Technology(电子科技大学; 武汉理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DiffSafeMerge是一种缓解扩散模型合并中后门继承的方法,通过评分源模块、收缩可疑贡献并在干净去噪损失预算下选择衰减,在多组实验中实现低攻击成功率和低FID。

AI 中文摘要

无条件扩散模型检查点合并假设源模型是良性的,但被篡改的公共检查点会在生成正常图像的同时传递潜伏的后门。由于无法获知被篡改的源、触发条件或目标,且广泛的净化可能会降低图像质量,因此缓解该问题十分困难。我们提出DiffSafeMerge(DSM),该方法使用小型未标记干净集和固定的、与攻击无关的压力探针对源模块进行评分,将可疑贡献向可信参考收缩,并在干净去噪损失预算下选择衰减系数。我们评估了4种攻击、2个数据集和21种目标条件:在14种源案例中,预期合并已在10种情况下实现最坏目标攻击成功率(ASR)为0;DSM保留了这些结果,且在剩余4种情况的3个随机种子中未记录到目标匹配,其中3种情况的基线ASR为48%至100%。在两个数据集上均实现最坏目标ASR为0的方法中,DSM在匹配的种子0比较中获得了最低的案例平均FID。

英文摘要

Unconditional diffusion checkpoint merging assumes benign sources, yet a compromised public checkpoint can transfer a dormant backdoor while clean generation appears normal. Mitigation is difficult without knowing the compromised source, trigger, or target, and broad sanitization may degrade image quality. We introduce DiffSafeMerge (DSM), which uses a small unlabeled clean set and fixed, attack-agnostic stress probes to score source blocks, shrink suspicious contributions toward a trusted reference, and select attenuation under a clean denoising-loss budget. We evaluate four attacks, two datasets, and 21 target conditions. Intended merging already has zero worst-target ASR in 10 of 14 source cases; DSM preserves these outcomes and records no target match in the remaining four over three seeds, including three with baseline ASR of 48--100\%. Among methods with zero worst-target ASR on both datasets, DSM obtains the lowest case-averaged FID in the matched seed-0 comparison.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑