arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25750cs.LGcs.CY

从权重中检测用于生成儿童性虐待材料的文本到图像的LoRA

Detecting CSAM Text-to-Image LoRAs From Weights

David Demitri Africa, Cate Heine, Nadine Staes-Polet, Kimberly Mai

AI总结:

研究如何从权重中检测用于生成CSAM的文本到图像的LoRA,利用LoRA更新的左上角奇异向量形成的指纹$u_1$,以人类主体年龄为代理,发现其能识别训练内容、跨模型泛化且对良性内容不判断,可直接从权重筛选有害LoRA。

AI中文摘要:

低秩适应(LoRA)微调使针对特定任务定制开放权重图像生成模型变得廉价且容易,这其中包括生成儿童性虐待材料(CSAM)。现有的审核依赖元数据或生成的输出,但元数据可能具有欺骗性,生成输出本身可能不可接受或非法。我们表明权重中存在更安全的信号。LoRA更新的左上角奇异向量形成了其最强学习变化的紧凑、无需推理的指纹($u_1$)。以人类主体年龄作为CSAM的良性代理,我们发现$u_1$能识别LoRA的训练内容,能跨基础模型泛化,并对不相关的良性内容不做判断。该信号对加权噪声、重新缩放和精度降低具有鲁棒性。这些结果表明,可以直接从有害LoRA的权重中进行筛选,而无需依赖元数据或生成有害输出。

英文摘要:

Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abuse material (CSAM). Existing moderation relies on metadata or generated outputs, but metadata can be deceptive and generating outputs may itself be unacceptable or illegal. We show that a safer signal lives in the weights. The top-left singular vectors of a LoRA's updates form a compact, inference-free fingerprint ($u_1$) of its strongest learned change. Using human-subject age as a benign proxy for CSAM, we find that $u_1$ identifies what a LoRA was trained on, generalizes across base models, and abstains on unrelated benign content. The signal is robust to additive weight noise, rescaling, and precision reduction. These results indicate that harmful LoRAs could be screened directly from their weights without relying on metadata or generating harmful outputs.

↑