发表机构
NAVER AI Lab(NAVER人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型两用知识的安全干预问题,提出“令牌接种”方法,通过绑定和分支操作,在保留有害知识的同时实现选择性拒绝,在有害领域降低准确率,保持良性领域性能,取得较好安全效用权衡,且拒绝选择性可控。
AI 中文摘要
对两用知识的安全干预通常在销毁有害内容(如遗忘、过滤)和在输出层抑制它(如拒绝训练)之间做出选择,这两种方法都会在相邻领域能力或过度拒绝方面付出代价。我们认为正确的操作是条件设定,而非减少:我们表明有害知识可以保留在模型中,并通过特权控制令牌进行行为控制。我们的方法“令牌接种”引入了一种绑定和分支方法。首先,在持续预训练期间,通过在两用文档旁边插入特殊令牌来标记有害内容,使模型将标记与有害领域的底层语义绑定。其次,在监督微调期间,教导模型在特殊令牌存在时正确回答有害查询,在其不存在时拒绝,从而实现选择性拒绝而不删除两用知识。在有害领域(如WMDP - Bio),令牌接种将准确率从79%降至18%,同时保留基础模型93%的良性领域性能(如MMLU),在1B - 14B模型规模上实现了与遗忘和拒绝调整基线相比最佳的安全效用权衡。我们还表明拒绝选择性可通过条件信号的质量控制,预训练期间特定领域的语义绑定对于条件行为推广到记忆触发器之外至关重要。我们的结果表明,安全对齐作为一个条件设定问题比遗忘问题更好:当敏感知识在受控访问下保留时,行为控制比销毁时更精确。
英文摘要
Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain competence or over-refusal. We argue that the right operation is conditioning, not reduction: we show that hazardous knowledge can be retained in the model and behaviorally gated by a privileged control token. Our method, Token Inoculation, introduces a binding-and-branching approach. First, during continued pre-training, we mark hazardous content by inserting a special token alongside dual-use documents, so the model binds the marker to the underlying semantics of the hazardous domain. Second, during supervised fine-tuning, we teach the model to answer hazardous queries correctly when the special token is present and to refuse them when it is absent, thereby enabling selective refusal without removing dual-use knowledge. On hazardous domain (e.g., WMDP-Bio), Token Inoculation reduces accuracy from 79% to 18% while retaining 93% of the base-model's benign-domain performance (e.g., MMLU), achieving the best safety-utility trade-off against unlearning and refusal-tuning baselines across 1B-14B model scales. We further show that refusal selectivity is controllable through the quality of the conditioning signal and that domain-specific semantic binding during pre-training is critical for the conditional behavior to generalize beyond memorized triggers. Our results suggest that safety alignment is better cast as a conditioning problem than a forgetting one: behavioral control is more precise when sensitive knowledge is retained under controlled access than when it is destroyed.
Comments23 pages, 13 figures, 8 tables