发表机构
University of Cambridge; University of California, Berkeley; Arcadia Impact; Center on Long-Term Risk(剑桥大学; 加州大学伯克利分校; Arcadia Impact; 长期风险中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对接种提示在共现非期望行为场景下的不足,提出分层接种提示(SIP),利用小干净子集过采样并接种其余样本,显著减少非期望行为且保留更多期望行为,并引入后门稀释与密码锁定接种以进一步控制触发。
AI 中文摘要
监督微调在教授语言模型期望行为的同时,也可能教会其非期望行为。接种提示(IP)旨在通过在训练期间请求非期望行为并在推理时移除该请求来限制非期望的泛化。然而,非期望行为仍可能在无关提示下出现。IP还可能阻碍期望行为的学习。我们在两种行为在大多数训练样本中共现的设置下解决这些局限性,因此过滤掉带有非期望行为的样本只会留下一个较小的干净子集。我们引入了分层接种提示(SIP)。SIP利用一个小的干净子集来证明期望行为应在不同情境下无需非期望行为而持续存在。SIP在多样化的非触发提示下对这些干净样本进行过采样,同时对其余样本进行接种。SIP显著减少了非期望行为的表达,同时比IP保留了更多期望行为。即使我们将IP扩展为以与SIP相同的速率对同一干净子集进行过采样,这些收益仍然存在。此外,SIP在我们测试的所有有害建议设置中均产生了更低的涌现性错位率。SIP还可以进一步扩展,即使在明确请求非期望行为的提示下也能限制该行为。我们引入了后门稀释,它削弱了接种提示下的表达,以及密码锁定接种,它将触发集中在指定的密码上。综合来看,我们的研究结果表明,改变一个小干净子集的训练情境可以显著改善选择性泛化。
英文摘要
Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompts. IP can also hinder learning of the desired behaviour. We address these limitations in settings where both behaviours co-occur in most training examples, so filtering out examples with undesired behaviour leaves only a small clean subset. We introduce stratified inoculation prompting (SIP). SIP leverages a small clean subset to demonstrate that desired behaviour should persist without the undesired one across different contexts. SIP oversamples these clean examples under diverse non-eliciting prompts while inoculating the rest. SIP substantially reduces expression of undesired behaviour while preserving more of the desired behaviour than IP. These gains persist even when we extend IP to oversample the same clean subset at the same rate as SIP. Moreover, SIP yields lower emergent misalignment rates in all harmful-advice setups we tested. SIP can be further extended to limit the undesired behaviour even under prompts that explicitly request it. We introduce backdoor dilution, which weakens expression under the inoculation prompt, and password-locked inoculation, which concentrates elicitation on a designated password. Taken together, our findings show that changing the training contexts for a small clean subset can significantly improve selective generalisation.