arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15886cs.CL

用学习到的新词进行中期接种训练

Inoculation Midtraining with Learned Neologisms

Kyle O'Brien, Edward James Young, Puria Radmard, Nathalie Kirch, Cameron Tice, Tomek Korbak, David Demitri Africa

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出中期接种训练,通过在中期训练中引入新词标记不安全上下文,减少后期训练中的错位,同时保留良性属性迁移,但效果未超标准方法且边界易泄漏。

中文摘要 AI 辅助

大型语言模型(LLMs)在后期训练中经常同时学习到理想和不理想的属性。我们研究中期训练(一个更早的训练阶段)是否能影响这些属性中哪些会在后续泛化。我们引入了“中期接种训练”(Inoculation Midtraining)技术,该技术教基础模型将不安全行为归属于一个指定的<quarantine_token>上下文,如中期训练期间引入的<quarantine_token>新词(一个新token)所示,然后在该上下文内对模型进行不安全数据的后期训练。随后,我们在上下文之外评估模型,系统提示中排除<quarantine_token>新词。在监督微调和强化学习后期训练机制中,我们发现中期接种训练可以减少错位,同时保留良性数据属性的迁移(例如,用德语或莎士比亚散文风格说话)。然而,我们的方法并未优于标准的“接种提示”(Inoculation Prompting),对训练配置敏感,并产生一个泄漏的边界,附近的上下文线索可以重新激活该边界。这些结果表明,通过中期训练引入的学习关联进行接种可以塑造选择性泛化。尽管如此,在方法成为开发者安全框架中的承重组件之前,还需要更多工作。

英文摘要

Large language models (LLMs) often learn both desirable and undesirable properties during post-training. We study whether midtraining, an earlier training stage, can shape which of these properties later generalise. We introduce Inoculation Midtraining, a technique that teaches a base model that unsafe behaviour belongs to a designated <quarantine_token> context, as indicated by the <quarantine_token> neologism (a new token) introduced during midtraining, and then post-trains the model on unsafe data within that context. We then evaluate the model outside the context, with the <quarantine_token> neologism excluded from the system prompt. Across supervised fine-tuning and reinforcement learning post-training regimes, we find that Inoculation Midtraining can reduce misalignment while preserving the transfer of benign data properties (e.g., speaking in German or Shakespearean prose). However, our approach does not outperform standard Inoculation Prompting, is sensitive to training configuration, and produces a leaky boundary that nearby contextual cues can reactivate. These results show that inoculation with a learned association introduced via midtraining can shape selective generalisation. Still, more work is needed before this approach can become a load-bearing component in a developer's safety framework.

发表机构

  • Geodesic Research(大地测量研究)
  • OpenAI
  • UK AI Security Institute(英国人工智能安全研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑