AI 中文总结
本研究发现安全防护模型存在拒绝线索捷径,插入拒绝线索可翻转其对有害内容的判定,经稀疏互补掩码干预后,拒绝线索引发的检测失败率相对降低约79%,同时保留标准检测性能。
AI 中文摘要
安全防护模型被广泛用于过滤有害内容,通常通过对标注的提示-响应对进行监督微调来训练。我们审计了两个广泛使用的安全防护训练数据集WildGuardMix和GR-Train,发现在对有害提示的响应中,拒绝表达几乎仅与无害标签共同出现。这种不平衡催生了我们所称的拒绝线索捷径:在有害响应中插入拒绝线索可将防护模型的判定从有害翻转为无害。该捷径不仅影响在这些数据集上训练的防护模型,还影响训练数据未公开的官方发布模型如LlamaGuard3和Qwen3Guard。它在响应的不同位置均存在,且在同一模型家族的较小变体中通常表现更强。为缓解该问题,我们采用稀疏互补掩码作为轻量级事后干预措施,无需重新训练即可识别并抑制一小部分与捷径相关的注意力头和MLP神经元。在两个主要基准上,该干预措施使拒绝线索引发的响应初始检测失败率相对降低约79%,同时保留了标准检测性能。尽管使用单个响应位置的线索进行优化,但抑制效果可迁移至未见过的位置和数据集,表明不同位置的捷径表现部分由共享内部组件介导。进一步分析提供证据表明,对捷径的依赖与对合法拒绝的识别在功能上部分可分离,因为抑制捷径能广泛保留防护模型识别真实拒绝的能力。
英文摘要
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard training datasets, WildGuardMix and GR-Train, and find that among responses to harmful prompts, refusal expressions co-occur almost exclusively with unharmful labels. This imbalance motivates what we term the refusal-cue shortcut: inserting a refusal cue into a harmful response could flip the guard's verdict from harmful to unharmful. The shortcut affects not only guards trained on these datasets but also officially released models such as LlamaGuard3 and Qwen3Guard whose training data is undisclosed. It persists across response positions and is generally stronger in smaller variants within a family. To mitigate it, we adapt sparse complementary masking as a lightweight post-hoc intervention that identifies and suppresses a small set of shortcut-associated attention heads and MLP neurons without retraining. On two primary benchmarks, the intervention achieves an approximately 79% relative reduction in response-initial detection failures induced by refusal cues, while preserving standard detection performance. Although optimized using cues at a single response position, the suppression effect transfers to unseen positions and datasets, suggesting that shortcut manifestations across positions are partly mediated by shared internal components. Further analysis provides evidence that shortcut reliance and legitimate refusal recognition are partially functionally separable, as suppressing the shortcut broadly preserves the guard's ability to recognize genuine refusals.
Comments13 pages, 2 figures, and 14 tables. Includes supplementary material in the appendix