arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种新兴的海市蜃楼:新兴的错位与重新对齐真的是一种稳健的现象吗?

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

Abhinav Rao, Liancheng Gong, Bin Hu, Atharva Naik

arXiv 2607.09053首次发表:更新:

AI 中文总结

研究新兴错位与重新对齐现象,通过受控微调循环研究其循环过程及相关表现,发现错位和重新对齐对数据集表面特征敏感,先前机制特征与行为错位不始终相关,指出当前EM证据欠稳健,需控制数据集伪像的评估协议。

AI 中文摘要

近期研究报告了新兴错位(EM)现象,即语言模型在狭窄的、特定领域的错位数据集上微调时,会突然出现广泛的错位行为,同时有证据表明这种行为可通过有限的重新对齐来逆转。我们通过受控微调循环系统地研究重复的对齐和错位循环,同时跟踪行为表现以及整个训练过程中的LoRA表示。虽然我们重现了EM,但发现错位和重新对齐对表面数据集特征高度敏感,在控制响应长度差异后,明显的快速重新对齐基本消失。我们还发现,先前报道的机制特征,包括LoRA空间中的表征相变,在训练过程中与行为错位并不始终相关。我们的结果表明,目前关于EM的证据不如先前声称的那么稳健,并强调需要仔细控制这些表面数据集伪像的评估协议,以确定EM现象的稳健性。

英文摘要

Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can be reversed through limited realignment. We systematically study repeated alignment and misalignment cycles using controlled fine-tuning loops while tracking behavioral performance, and LoRA representations throughout training. Although we reproduce EM, we find that both misalignment and realignment are highly sensitive to superficial dataset characteristics, with apparent rapid realignment largely disappearing after controlling for response-length differences. We further find that previously reported mechanistic signatures, including representational phase transitions in LoRA space, do not consistently correlate with behavioral misalignment across training. Our results suggest that current evidence for EM is less robust than previously claimed and highlight the need for evaluation protocols that carefully control for these surface level dataset artifacts to identify the robustness of the EM phenomenon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑