arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22215cs.CLcs.LG

论大语言模型中潜意识学习的缓解

On Mitigation of Subliminal Learning in Large Language Models

  • Algoverse AI Research(Algoverse AI 研究院)

机构由 AI 辅助整理,请以论文原文为准。

Atsushi Yanagisawa, Brendan Gho, Rajendran Ramesh Babu Manoj Narender, Kevin Zhu, Madhur Panwar, Antonio Mari

AI总结:

本文研究大语言模型微调中的潜意识学习现象,提出退火KL正则化的阈限训练方法,有效抑制特征习得并保留任务性能,优于现有缓解策略。

AI中文摘要:

知识蒸馏可以将教师模型中非预期的行为特征通过看似与该特征语义无关的训练数据传递给学生模型,这一现象被称为潜意识学习。尽管近期研究已证实该效应,但其训练动态及缓解方法仍未得到充分探索。我们在参数量从1.5B到8B的开源权重语言模型中研究潜意识学习,涵盖Qwen、Gemma和Llama系列,在数字序列和思维链场景下进行实验。我们不仅评估最终模型,还在微调过程中跟踪特征相关概率,发现潜意识习得可能高度非单调,存在瞬时尖峰、反转以及特征特定的迁移失败。随后,我们引入了阈限训练(liminal training),一种退火KL正则化的微调方法,用于约束模型早期偏离基础模型。在我们的实验中,阈限训练显著减少了潜意识特征的习得,同时很大程度上保留了任务收益,优于释义和层冻结等缓解策略。该效应也超越了动物偏好:在法语回答风格实验中,阈限训练抑制了语言迁移,同时保留了GSM8K的大部分改进。最后,我们表明KL时机至关重要:早期正则化比晚期正则化更有效,并且扫描正则化强度揭示了任务学习与特征抑制之间的经验性权衡。

英文摘要:

Knowledge distillation can transmit unintended behavioral traits from a teacher model to a student through training data that appear semantically unrelated to those traits, a phenomenon known as subliminal learning. Although recent work has established this effect, its training dynamics and mitigation remain underexplored. We study subliminal learning in open-weight language models ranging from 1.5B to 8B parameters, covering the Qwen, Gemma, and Llama families in number-sequence and chain-of-thought settings. Rather than evaluating only final models, we track trait-related probabilities throughout fine-tuning and find that subliminal acquisition can be highly non-monotonic, with transient spikes, reversals, and trait-specific failures of transfer. We then introduce liminal training, an annealed KL-regularized fine-tuning method that constrains early drift from the base model. Across our experiments, liminal training substantially reduces subliminal trait acquisition while largely preserving task gains, outperforming paraphrasing and layer freezing as mitigation strategies. The effect also extends beyond animal preferences: in a French-language response-style experiment, liminal training suppresses language transfer while retaining much of the GSM8K improvement. Finally, we show that KL timing matters: early regularization is more effective than late regularization, and sweeping the regularization strength reveals an empirical trade-off between task learning and trait suppression.

补充信息

↑