arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

复现与评估开放权重模型中潜意识学习的可泛化性

Reproducing and Evaluating the Generalizability of Subliminal Learning in Open-Weight Models

Daan van der Weijden, Nathan Brack, Selene Baez Santamaria

arXiv 2609.12586首次发表:更新:

发表机构

University of Zurich(苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本论文复现并扩展了开放权重模型中潜意识学习的研究,验证了原始主张,但发现传递效果因特征、任务和模型而异,并非普遍适用。

AI 中文摘要

在这篇复现论文中,我们研究了潜意识学习,这是蒸馏的一种后果,即教师模型通过语义无关的数据传递行为偏好特征。原始论文探讨了两种类型的特征(动物偏好和错位)、三种数据模态(数字序列、代码和思维链)以及多个模型家族。我们复现了他们的实验,并沿三个方向扩展了设置:新的偏好类别(演员和政治家)、新任务(国际象棋走子生成)以及额外的开放权重模型(Ministral8B)。我们还对数字任务的答案空间大小(1位、2位和3位数字序列)进行了受控消融实验。我们专注于在HuggingFace上具有可访问检查点的开放权重模型,因为原始论文的GPT-4.x微调已不再可用。我们的复现支持原始论文的主张,但我们的扩展表明这些主张并非普遍适用,因为传递强度因特征和任务而异,且一个模型几乎没有任何效果。

英文摘要

In this reproduction paper we investigate subliminal learning, a consequence of distillation where teacher models transmit behavioral preference traits through semantically unrelated data. The original paper explores two types of traits (animal preferences and misalignment), three data modalities (number sequences, code, and chain of thought), and several model families. We reproduce their experiments and extend the setup along three axes: new preference categories (actors and politicians), a new task (chess move generation), and an additional open-weight model (Ministral8B). We also run a controlled ablation on the numbers task's answer-space size (1-, 2-, and 3-digit sequences). We focus on open-weight models with accessible checkpoints on HuggingFace, since the original paper's GPT-4.x fine-tuning is no longer available. Our reproduction supports the original paper's claims, but our extensions show they are not universal as transmission strength varies across traits and tasks, and one model shows almost no effect at all.

CommentsAccepted at BlackBoxNLP@EMNLP'26 (The 9th BlackboxNLP Workshop Special Track: Reproducibility and Reliability in Interpretability Analyses)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑