通过中间干净图像估计实现安全文本引导图像生成的测试时缩放
Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
浏览论文内容
中文总结 AI 辅助
该研究针对文本到图像扩散模型的安全问题,提出利用中间干净图像估计和稀疏边际目标的测试时缩放方法,在 Stable Diffusion 上实现了更优的安全性能。
中文摘要 AI 辅助
确保文本到图像扩散模型的安全性和政策合规性仍是一项关键挑战,因为良性或对抗性提示词常能生成禁止内容,例如裸体和受保护知识产权。虽然基于训练的遗忘方法有效,但计算成本高昂且易对通用能力造成灾难性干扰。相反,现有的测试时防御主要以提示词为中心,仅依赖修改文本描述,而忽略视觉信号用于检测。在本文中,我们提出利用生成过程中估计的中间干净图像,并采用稀疏边际目标来检测禁止概念。当检测到违规时,我们通过截断反向传播在文本条件空间中优化结构化低秩残差来立即干预。该设计支持权重保留检测,随着最大预算增加,非违规推理延迟几乎保持不变,并通过测试时缩放提供安全性能的灵活性。在 Stable Diffusion v1.4 和 v3.5 上针对裸体去除、知识产权保护和风格擦除的广泛实验表明,与现有权重保留基线相比,在抑制、保真度和保留方面表现出优越性能,为安全生成部署提供了可扩展且灵活的解决方案。
英文摘要
Ensuring safety and policy compliance in text-to-image diffusion models remains a critical challenge, as benign or adversarial prompts can often elicit prohibited content, e.g. nudity and protected intellectual property. While training-based unlearning methods are effective, they are computationally expensive and prone to catastrophic interference with general capabilities. Conversely, existing test-time defenses are primarily prompt-centric, relying on modifying textual descriptions only, and overlook the visual signals for detection. In this paper, we propose to leverage the intermediate clean image estimated during the generation process and employ a sparse margin objective to detect prohibited concepts. When a violation is detected, we immediately intervene by optimizing a structured low-rank residual in the text-conditioning space via truncated backpropagation. This design allows weight-preserving detection, keeps non-violating inference latency nearly unchanged as the maximum budget increases, and offers flexibility in safety performance via test-time scaling. Extensive experiments on Stable Diffusion v1.4 and v3.5 across nudity removal, IP protection, and style erasure demonstrate superior performance across suppression, fidelity and preservation compared to prior weight-preserving baselines, providing a scalable and flexible solution for safe generative deployment.