新兴的不一致性招募了预先存在的角色子空间
Emergent Misalignment Recruits a Pre-existing Persona Subspace
查看机构详情
- Qwen Team(通义团队)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究新兴不一致现象,通过从冻结模型提取角色子空间发现其在微调前就存在,微调第一步提升不一致幅度,投影子空间可防广泛不一致,注入子空间会致模型不一致,多种编辑无效,传播不良数据增加不一致。
中文摘要 AI 辅助
在少量不良建议上微调对齐的语言模型,会使其在与训练数据无关的问题上产生广泛的不一致,即新兴不一致现象。我们探究为何这种少量训练能产生泛化效果,发现少量微调会招募微调前模型中就存在的角色结构。通过对比教师强制从冻结的指令微调模型中提取每个领域的角色子空间,发现4个不相关领域在657倍于随机子空间零空间的情况下共享一个低秩核心,且该核心的82%位于匹配多样性构建的风格核心之外。在不安全代码上微调的第一步比将相同代码视为教育性代码时更难提升广泛不一致的幅度,并预测到375步时的幅度变化。在微调过程中从残差流中投影出子空间可防止广泛不一致(判断生成中从27.7%降至0.0%),而匹配秩的随机子空间则无作用;将其注入未微调模型会导致不一致性随剂量增加至45.4%,超过与之对比的微调模型。应用于权重梯度的相同投影无效,三种事后权重编辑也无法改变这种情况:最剧烈的编辑抑制了行为而非消除它,且被消融的结构在编辑清除的子空间内重新形成。在4个领域传播固定预算的不良数据产生的广泛不一致比机械权重叠加和匹配多样性共同作用的结果更多。所有测量均来自一个14B模型;提取来自对齐的指令微调检查点,这使得结构的来源不明;防止不一致的干预也消除了少量训练的行为。
英文摘要
Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment. We ask why the narrow lesson generalizes at all, and we find that narrow fine-tuning recruits a persona structure that is present in the model before the fine-tune exists. From a frozen instruction-tuned model (Qwen2.5-14B-Instruct) we extract per-domain persona subspaces by contrastive teacher forcing and find that 4 unrelated domains share one low-rank core at 657x a random-subspace null, with 82% of that core lying outside a style core built at matched diversity. The literal first optimizer step of fine-tuning on insecure code climbs a broad-misalignment margin harder than the same code framed as educational, and forecasts realized margin movement out to 375 steps. Projecting the subspace out of the residual stream throughout fine-tuning prevents broad misalignment (27.7% to 0.0% of judged generations) while a matched-rank random subspace changes nothing; injecting it into the never-fine-tuned model induces misalignment that grows with dose to 45.4%, past the fine-tuned model it is measured against. The same projection applied to the weight gradient is inert, and three post-hoc weight edits leave the disposition in place: the sharpest edit suppresses the behavior rather than removing it, and the ablated structure re-forms inside the subspace the edit cleared. Spreading a fixed budget of bad data across 4 domains produces more broad misalignment than mechanical weight superposition and matched diversity jointly account for. All measurements come from one model at 14B; the extraction is from an aligned instruction-tuned checkpoint, which leaves the structure's provenance open; and the intervention that prevents misalignment also abolishes the narrow trained behavior.