arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11025cs.CL

基于角色特征的涌现失配数据归因

Data Attribution of Emergent Misalignment with Persona Features

  • Bonn-Aachen International Center for Information Technology, University of Bonn(波恩-亚琛信息技术中心,波恩大学)
  • Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Clemens Vetter, David Kaczér, Lucie Flek, Florian Mai

AI总结:

本研究基于SAE分析揭示EM源于角色特征,操控特征可双向控制EM,发现人类文本无法可靠诱导EM而合成指令响应对可以,明确响应结构或模型措辞的关键作用。

AI中文摘要:

涌现失配(EM)是指在狭窄任务上微调语言模型会导致其在不相关领域出现有害行为的现象。主流机制解释将EM归因于角色特征:预训练过程中获得的潜在方向,失配微调会放大这些特征。本研究探究这些特征的来源:哪些预训练文档会激活它们,以及自然存在的人类编写文本是否足以诱导EM。通过基于稀疏自编码器(SAE)的模型差异分析,在四个开放权重模型上的研究发现,越狱角色、讽刺、欺骗和操纵相关的特征会被失配微调放大,而与安全相关和助手身份相关的特征则被抑制。操控单个特征可双向控制EM:它能在对齐模型中诱导最高达62%的失配率,超过失配微调本身达到的35%,还能将失配模型重新对齐至接近基线的失配率。将因果特征归因于包含100万份预训练网页文档的语料库后,检索到与反派角色、支配和有害能动性相关的语义相关叙事。然而,即使将这些人类编写文档重新格式化为助手风格的响应,对其进行微调也无法可靠地诱导EM,而源自相同内容的合成指令-响应对则可以做到,且能跨模型家族迁移。因此,仅语义相关性是不够的:响应结构或模型生成的措辞在诱导EM中发挥重要作用。

英文摘要:

Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diffing across four open-weight models, we find that features related to jailbreak personas, sarcasm, deception, and manipulation are amplified by misalignment fine-tuning, while safety-relevant and assistant-identity features are suppressed. Steering individual features controls EM in both directions: it induces misalignment rates of up to 62% in aligned models -- exceeding the 35% reached by misalignment fine-tuning itself -- and re-aligns misaligned models to near-baseline misalignment rates. Attributing the causal features to a corpus of one million pre-training web documents retrieves semantically relevant narratives about villainous characters, domination, and harmful agency. However, fine-tuning on these human-written documents does not reliably induce EM, even after reformatting into assistant-style responses, whereas synthetic instruction-response pairs derived from the same content do -- and transfer across model families. Semantic relevance alone is therefore not sufficient: response structure or model-generated phrasing plays an important role in inducing EM.

补充信息

↑