arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29118cs.AIcs.CLcs.LG

涌现失配并非神奇现象

Emergent Misalignment Is Not Magical

  • The University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

Mingxuan Li, Qirun Dai, Heran Wang, Chenhao Tan

AI总结:

该研究揭示涌现失配(EM)是可预测的数据依赖泛化现象,通过表征距离可预测其效果,还扩展了泛化指标,能可靠预测EM模型在语义保留提示扰动下的邪恶程度。

AI中文摘要:

在窄范围有害数据集上微调大型语言模型(LLMs)会导致广泛的失配,这一现象被称为涌现失配(EM)。EM对AI安全及我们对LLMs的理解构成挑战,过往研究常将EM视为意外行为,通过泛化失配方向或拟人化为获得“邪恶人格”来解释,但这些解释背后的机制仍不明确。本研究表明,EM是可预测且依赖数据的泛化现象。通过检查基础模型对EM训练数据和评估提示的表征,我们发现EM训练后的“邪恶程度”可通过表征距离高度预测:评估提示与训练数据质心越接近,EM模型对其引发的“邪恶程度”越高(在12种模型-数据集设置下,平均斯皮尔曼相关系数为-0.73)。基于此分析,我们进一步揭示EM的本质:(1)其效果随训练数据格式显著变化;(2)不存在可在不同EM模型间迁移的通用失配方向;(3)EM的效果与人格变化根本不同。此外,我们将EM泛化指标从标量距离扩展为特定数据集的泛化方向,该指标能可靠预测EM模型在语义保留的提示扰动(包括附加随机标记和释义)下的“邪恶程度”,而其他方法无法实现可靠泛化。

英文摘要:

Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM). EM poses a challenge for AI safety and our understanding of LLMs. Prior work often frames EM as an unexpected behavior, and explains it by appealing to general misalignment directions or anthropomorphizing it as acquiring an evil persona. However, the mechanisms behind these framings remain obscure. In this work, we show that EM is a predictable and data-dependent generalization phenomenon. By examining the base model's representation of EM training data and evaluation prompts, we find that evilness after EM training is highly predictable from representational distance: the closer an evaluation prompt is to training data centroid, the more evilness it elicits from EM models after training (with an average Spearman correlation of -0.73 across 12 model-dataset settings). Building upon this analysis, we further demystify EM by showing that (1) its effectiveness changes significantly based on training data format; (2) there is not a general misalignment direction that transfers across different EM models; (3) the effect of EM is fundamentally different from persona changes. Furthermore, we extend the EM generalization metric from a scalar distance to a dataset-specific generalization direction, which robustly predicts EM models' evilness under semantics-preserving prompt perturbations including appending random tokens and paraphrasing, where other methods do not reliably generalize.

↑