发表机构
Alibaba AAIG(阿里巴巴AAIG)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出InGuard内部护栏框架,在模型表示上分级提示、修改嵌入并中途检测,实现高安全率与低干扰,显著减少参数和计算开销。
AI 中文摘要
现代文生图(T2I)模型能够根据任意用户提示生成高质量图像,但也同样容易生成不宜工作场所(NSFW)的内容。传统的外部护栏由两个组件构成:一个在生成之前检查风险的提示分类器,以及一个在图像完全生成后检查图像的后期图像分类器。在这种设计中,两个分类器都在生成流程之外运行,不使用模型自身的表示。这种分离可能限制提示筛选的准确性,而图像侧检查仅在全部生成成本已花费之后才运行。此外,被标记的提示只能被拒绝,即使它可以被调整以生成安全图像。在这项工作中,我们提出了内部护栏(InGuard),一种在流程内部基于模型自身表示运行的安全框架,不改变基础模型参数。首先,一个风险分类器基于文本编码器的嵌入将每个提示分级为不安全、有风险或良性,无需外部语言模型。其次,SAGE(用于嵌入的软门控非对称护栏)修改有风险提示的嵌入,旨在返回安全图像而非拒绝。第三,一个潜在检测器在去噪中途检查单步干净潜在估计,达到接近图像级的性能,并在检测到风险时停止生成。我们还构建了RevGen安全基准,以在现实条件下评估T2I安全性:通过真实图像反向生成构建的10,000个提示,并带有重写步骤以提供受控的知识产权(IP)角色,涵盖分级色情/血腥风险、分类IP风险和良性负面案例。在五个开放权重T2I模型上,InGuard达到97.9-98.8%的安全率,匹配或超过外部护栏,良性干扰减少57.5-73.5%,参数减少约3.7倍,并跳过50-55.6%的去噪步骤。
英文摘要
Modern text-to-image (T2I) models generate high-quality images from arbitrary user prompts, yet they can just as easily produce not-safe-for-work (NSFW) content. Conventional outer guardrails consist of two components: a prompt classifier that checks for risk before generation, and a post-hoc image classifier that checks the fully generated image. In this design, both classifiers operate outside the generation pipeline and do not use the model's own representations. This separation can limit prompt-screening accuracy, while the image-side check runs only after the full generation cost has been spent. Moreover, a flagged prompt can only be rejected, even when it could be adjusted to produce a safe image. In this work, we propose the Inner Guardrail (InGuard), a safety framework that works inside the pipeline on the model's own representations, leaving base-model parameters untouched. First, a risk classifier grades each prompt as unsafe, risky, or benign based on the text encoder's embeddings, with no external language model. Second, SAGE (Soft-gated Asymmetric Guardrail for Embeddings) modifies the embeddings of risky prompts, aiming to return a safe image instead of a refusal. Third, a latent detector checks the one-step clean latent estimate midway through denoising, reaching nearly image-level performance and halting generation when risk is detected. We also construct the RevGen Safety Benchmark to evaluate T2I safety under realistic conditions: 10,000 prompts built through real-image reverse generation, with a rewriting step that supplies controlled intellectual-property (IP) characters, covering graded porn/gore risks, categorical IP risks, and benign negatives. Across five open-weight T2I models, InGuard reaches 97.9-98.8% safety rate, matching or exceeding the outer guardrail, with 57.5-73.5% less benign disturbance, ~3.7x fewer parameters, and 50-55.6% of denoising steps skipped.