arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33634cs.CLcs.AI

安全重构:通过掩码扩散进行生成式建模构建强大的安全护栏

Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

  • University of Neuchâtel(纳沙泰尔大学)
  • Delft University of Technology(代尔夫特理工大学)
  • IBM Research(IBM研究)
  • University of Turin(都灵大学)

机构由 AI 辅助整理,请以论文原文为准。

Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Chen

AI总结:

针对现有安全护栏模型训练目标狭窄、依赖捷径特征的问题,本文提出LLaDA-Guard,利用掩码扩散语言模型进行类条件重建,将监督扩展到每个令牌,在多个基准上提升性能与校准,并支持令牌级风险定位与安全重写。

AI中文摘要:

护栏模型是语言模型与有害输出之间的最后一道防线,然而其训练目标却出奇地狭窄。现有的护栏模型学习从对话上下文中预测一个单一的判定令牌,将监督集中在单个目标上。这导致了结构性问题:模型会依赖捷径特征,过度自信,并且对安全证据在序列中出现的位置敏感,而非其在完整上下文中的作用。我们提出了一种不同的框架。我们的LLaDA-Guard不是从文本中预测标签,而是询问哪个标签能更好地解释文本:在每个标签假设下对提示或响应进行评分,并根据其差异进行分类。这将监督转移到被审核区域的每个令牌上,迫使模型考虑完整内容而非最具区分性的片段。我们使用掩码扩散语言模型实例化了这一想法,通过LoRA使用类条件重建目标对LLaDA-8B-Instruct进行微调,且除了基础模型外无需任何架构更改。LLaDA-Guard在七个保留的安全基准上,相对于在更强骨干网络上训练的判别式基线,平均排名领先,同时表现出显著更好的置信度校准(ECE 0.0875对比Qwen3Guard的0.1384),对带有不安全线索的良性提示的过度防御更少,并且在审核响应时提示泄漏更少。其生成特性还使得令牌级风险定位成为自然的副产品,产生了一个无需额外训练即可将不安全提示重写为安全等效提示的流程,实现了60.7%的平均转换到安全率。

英文摘要:

Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA-Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA-Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with unsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting unsafe prompts into safe equivalents without additional training and achieving a 60.7% average conversion-to-safe rate.

↑