阈下学习(Subliminal Learning, SL)是非语义知识蒸馏
Subliminal Learning is Non-Semantic Distillation
AI总结:
该研究探究阈下学习(SL)的机制,发现非语义权重结构是关键,引导向量可生成阈下数据,相关梯度或助力数据审计,凸显识别训练数据潜在信号的重要性。
AI中文摘要:
阈下学习(Subliminal Learning, SL)是现代语言模型展现出的一种令人惊讶的泛化类型,它允许通过从教师模型看似不相关或随机的合成数据中进行蒸馏,将教师模型的偏差或行为迁移到学生模型。这给确保AI系统保持可预测性并进行安全训练带来了挑战,因为对输入数据的标准审计无法捕捉到隐藏的阈下信号。在此,我们研究了关于阈下学习的促成机制和驱动因素的几个开放问题。首先是偏差在数据中编码的过程的本质。我们发现,通过向教师模型和学生模型的权重添加高斯噪声,在Gemma模型中阈下迁移的幅度增加了1.9倍,在Llama模型中增加了1.3倍,这表明非语义权重结构起着关键作用。我们证明,除了之前研究中使用的提示和微调之外,还可以将引导向量(steering vectors)应用于教师模型以生成阈下数据。对在引导数据和提示数据上训练的学生模型的激活进行的分析表明,学生不仅继承了教师偏差的语义含义,还继承了用于应用该偏差的干预类型:经引导的学生模仿引导向量,而经提示的学生则不会。此外,引导型阈下数据的梯度与教师的引导向量呈线性相关,这为数据审计带来了希望。更广泛地说,随着合成数据成为前沿训练流程的核心,能够识别训练数据中隐藏的潜在信号变得至关重要。
英文摘要:
Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a teacher model to a student by distilling from seemingly unrelated or random synthetic data from the teacher. This presents challenges in ensuring AI systems remain predictable and are trained safely, as standard auditing of the input data would not catch the hidden subliminal signal. Here, we investigate several open questions as to the enabling mechanisms and drivers of SL. First is the nature of the process by which biases are encoded in the data. We find that by adding Gaussian noise to the weights of the teacher and student models, the magnitude of subliminal transfer is increased by a factor of 1.9 in Gemma and 1.3 in Llama, suggesting that non-semantic weight structures play a crucial role. We show that steering vectors can be applied to the teacher to produce subliminal data, in addition to prompting and finetuning as used in previous studies. Analysis of the activations of the student models that have been trained on steered and prompted data demonstrates that students inherit not just the semantic meaning of the teacher's bias, but also the type of intervention that was used to apply it: steered students imitate steering vectors, prompted students do not. Additionally, the gradients of steered subliminal data show a linear correlation with the teacher's steering vectors, showing promise for data auditing. More broadly, as synthetic data becomes central to frontier training pipelines, being able to see the latent signals hidden in training data becomes paramount.