数据归因能否过滤潜意识学习?并不可靠
Can Data Attribution Filter Out Subliminal Learning? Not Reliably
- Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所)
- International Institute of Information Technology Hyderabad(海得拉巴国际信息技术学院)
- Technische Universität Berlin(柏林工业大学)
- Zuse School ELIZA(楚泽ELIZA学院)
- Technological University Dublin(都柏林理工大学)
- BIFOLD – Berlin Institute for the Foundations of Learning and Data(BIFOLD – 柏林学习与数据基础研究所)
- EleutherAI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究评估三种基于梯度的数据归因方法在过滤语言模型潜意识学习中的可靠性,发现EK-FAC在标记级过滤中有效,但整体成功率不一致,不如分歧标记基线。
AI中文摘要:
潜意识学习使语言模型能够通过训练数据传递与这些行为特征无明显语义关联的行为特征,从而削弱了基于内容的数据过滤作为安全干预措施的有效性。训练数据归因提供了一种替代方案:它识别出导致特定模型行为的训练样本,独立于其语义内容,因此可能恰好适用于语义检查失效的情况。我们在三个模型上评估了三种基于梯度的归因方法(GradCos、一种对比性GradCos变体以及EK-FAC),并将它们与分歧标记(一种先前被证明能够定位潜意识学习的强基线,尽管它需要访问反事实教师模型)进行比较。在标记级别进行过滤时,EK-FAC缓解了该效应的显著部分,其他方法提供的益处甚微,且所有方法大多不如分歧标记。对每个方法而言,过滤整个样本的效果较差,尽管在此设置中EK-FAC通常比分歧标记提供更强的信号。不同方法和设置下的成功并不一致:对某些模型-偏好组合有效的变体在其他组合中失效,我们未能为这些差异找到一致的解释。我们的结果表明,基于梯度的归因在某些设置中可以识别导致潜意识学习的数据,但某些近似方法比其他方法更可靠。
英文摘要:
Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it identifies the training examples responsible for a given model behavior, independent of their semantic content, and so may apply in exactly the cases where semantic inspection fails. We evaluate three gradient-based attribution methods (GradCos, a contrastive GradCos variant, and EK-FAC) across three models, comparing them against divergence tokens, a strong baseline previously shown to localize subliminal learning (albeit one that requires access to counterfactual teacher models). Filtering at the token level, EK-FAC mitigates a significant part of the effect, the other methods provide little benefit, and all mostly fall short of divergence tokens. Filtering entire samples is less effective for every method, though EK-FAC often gives a stronger signal than divergence tokens in this setting. Success is inconsistent across methods and settings: variants that work well for some model-preference combinations fail for others, and we do not identify a consistent explanation for these differences. Our results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.