arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21869cs.CVcs.AI

GuardPaint:文本到图像生成的推测性安全解码

GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration

Shreyash Dhoot, Paras Dhiman, Arsh Abbas Naqvi, Aranbi Dutta, Aman Chadha, Vinija Jain, Amitava Das

首次发表
浏览论文内容

中文总结 AI 辅助

GuardPaint是一种无需修改基础模型的文本到图像生成安全防护框架,可在扩散轨迹内干预,通过轻量级审计器定位并修复不安全区域,在多类越狱提示词和主流T2I模型上显著降低有害内容生成且对图像质量影响极小。

中文摘要 AI 辅助

文本到图像(T2I)扩散模型具备强大的视觉生成能力,但其可控性带来了关键的安全挑战:对抗性提示词可引导去噪轨迹生成违反安全策略的内容,如露骨的裸体内容或暴力图像。现有安全防护措施大多仅在生成前通过提示词过滤或生成后通过图像分类发挥作用,未对扩散过程本身进行防护,且通常仅会拒绝生成而非对图像进行安全修复。本文提出GuardPaint,一种用于安全T2I生成的推测性解码框架,无需修改基础模型即可在扩散轨迹内部进行干预。一个轻量级审计器监控中间图像,定位不安全区域,并仅在必要时触发针对性的修复性图像修复;修复候选由策略对齐的图像修复模型生成,并通过受保护的锦标赛机制进行筛选,仅在编辑内容提升策略合规性且同时保留提示词保真度和感知质量时才接受该编辑。在SneakPrompt、MMA、PGJ、DACA和RABell这5类越狱提示词家族,以及UNet/流匹配模型(包括SD~1.5、SDXL、SD~3.5和FLUX.1-dev)上的实验表明,GuardPaint可降低攻击成功率和有害内容生成,同时对图像质量、提示词保真度及良性行为的影响极小。内容警告:本文包含涉及裸体和暴力的示例,部分读者可能会感到不适、痛苦或反感。

英文摘要

Text-to-image (T2I) diffusion models offer powerful visual generation, but their controllability creates a critical safety challenge: adversarial prompts can steer the denoising trajectory toward policy-violating content such as explicit nudity or graphic violence. Existing safeguards mostly act before generation through prompt filtering or after generation through image classification, leaving the diffusion process itself unguarded and often yielding only refusal rather than safe visual repair. We introduce GuardPaint, a speculative decoding framework for safe T2I generation that intervenes inside the diffusion trajectory without modifying the base model. A lightweight auditor monitors intermediate images, localizes unsafe regions, and triggers surgical inpainting repair only where needed. Candidate repairs are generated by a policy-aligned inpainter and selected through a guarded tournament that accepts edits only when they improve policy compliance while preserving prompt fidelity and perceptual quality. Across five jailbreak families SneakPrompt, MMA, PGJ, DACA, and RABell and UNet/flow-matching models including SD~1.5, SDXL, SD~3.5, and FLUX.1-dev. GuardPaint reduces attack success and harmful generations with minimal degradation to image quality, prompt fidelity, and benign behavior. Content warning: This paper contains examples involving nudity and violence that some readers may find disturbing, distressing, or offensive.

↑