arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37914cs.CLcs.AI

坏建议的不平等影响:利用训练数据归因调节涌现性错位

The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment

Gonçalo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman

首次发表
浏览论文内容

中文总结 AI 辅助

本研究利用训练数据归因量化有害示例对涌现性错位的影响,发现基于归因分数过滤数据可显著增强或减弱错位,且影响力分数在同模型过滤时效果最佳。

中文摘要 AI 辅助

在狭窄、错位的任务上微调大型语言模型可能会破坏其训练后的对齐,并诱发新颖的错位行为——这一现象被称为“涌现性错位”(EM)。EM 已与类似人格的表示相关联,其中微调可能通过放大有害或“邪恶”的人格来降低损失。目前尚不清楚训练数据的哪些属性驱动了这种效应:是否所有有害示例对错位的贡献大致相等,以及不同模型是否受到相同微调示例的同等影响。在本工作中,我们使用训练数据归因来定量估计每个有害示例对 EM 的贡献程度。我们通过重训练来评估归因的质量——一个合理的归因分数应使我们能够通过基于该分数过滤数据来增强或减弱 EM。基于分数的过滤可以显著增强或减弱 EM;我们发现,数据归因分数和黑盒有害性分数都能识别出关键示例。我们测试的所有模型在相同数据集上训练时都会变得错位,而影响力分数在过滤来自计算它们的同一模型的数据时表现最佳。我们发现,从我们测试的三个模型家族中得出的影响力分数具有跨模型泛化性,但这种泛化并不能恢复同模型过滤的性能。

英文摘要

Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining -- a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.

发表机构

  • EleutherAI

机构由 AI 辅助整理,请以论文原文为准。

↑