arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

移植、反转和防止错位角色:Qwen2.5中方法条件下的突发错位

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5

Lyndon Drake, Zandi Eberstadt

arXiv 2607.04510首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究Qwen2.5模型中突发错位,通过移植潜在角色方向诱导错位,消融自身方向可减少诱导,发现招募角色与方法和能力有关,还可筛查和防止其招募。

AI 中文摘要

突发错位(EM)在Qwen2.5模型中由潜在角色方向介导,移植该方向会引发广泛EM,消融自身方向可减少诱导。移植可作为测量方法,招募角色取决于方法和能力,其因果作用有条件,可筛查和防止招募,结果是对一个模型家族的控制案例研究。

英文摘要

Emergent misalignment (EM) --- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data --- is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights. Transplanting it into a model that shares only pretraining with its source induces broad EM (2.83 $\pm$ 0.26\% misaligned against a random-direction floor of $\sim$1.1\%), and ablating a model's own direction roughly halves an overt inducer's broadcast (21\% to 10\%). The transplant doubles as a measurement method, causally assaying directions that a source model represents but cannot itself express. Whether a fine-tune recruits this persona depends on method and capacity, and since low-rank PEFT is the cheaper regime at scale, the recruiting method is also the economical one. On Qwen2.5-32B, LoRA at low ranks on insecure code recruits it (3.4\% misaligned) while full SFT on identical data does not (0.3\%) and moves against the persona axis (drift--persona cosine $+0.17$ at rank 1 to $-0.10$), the far-inducer, high-capacity exception consistent with a representational-distance $\times$ capacity account. The persona's causal role is itself conditional. Steering a bad-medical SFT run away from the direction during training raises the broadcast from ${\sim}24\%$ to ${\sim}50\%$ while matched random controls stay at or below baseline, replicated across three training seeds, so removing the direction is no blanket recipe. Because recruitment is a loss-reducing shortcut that capacity renders redundant, it can be screened for and prevented in the tested instances. Persona loss-relevance at the SFT solution orders four inducers' broadcasts rank-perfectly within Qwen2.5, inoculation removes recruitment selectively (4.75\% to 0.0\%, code coherence 65\% to 87\%), and fine-tuning orthogonal to the single behaviour-derived axis reduces it persona-specifically.

Comments34 pages, 18 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑