arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

纠正而非删除:通过纠正性监督缓解突发性错位

Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision

Jacob Epifano

arXiv 2609.37624首次发表:更新:

AI 中文总结

本研究针对语言模型微调中的突发性错位问题,提出用纠正有害训练数据替代删除,实验表明纠正可将错位率降低约三分之一,优于删除方法。

AI 中文摘要

在狭窄的有害示范集(如糟糕的医疗建议)上微调语言模型,可能会使其在与无关问题上广泛错位,这一现象被称为突发性错位(EM)。通常的防御方法是找出违规行并删除它们,但行定位器在我们留出的测试中失败,且删除行的帮助效果低于预期。我们提出了一个不同的问题:给定一组固定的中毒行,纠正它们是否比删除它们更好?我们在混合了糟糕医疗建议和良性聊天数据的数据集上微调Qwen2.5-14B-Instruct,预先选择四分之一的毒行,然后要么删除它们,要么将每一行替换为针对同一提示的纠正答案,其余保持不变。替换这些行将EM率降低了约三分之一,并改善了留出医疗问题上的答案,而删除相同的行几乎没有可测量的效果。当一半毒行被纠正时,优势更大,并且在第二个基础模型和第二个错位模型实体上也成立。替换内容似乎很重要:对行进行改写但保留其糟糕建议没有明显益处,而随数据集分发的正确答案的表现与我们的改写器相当。已知通过对已中毒模型进行进一步微调可以重新对齐,但哪种数据起作用尚未直接比较。我们发现,对纠正数据进行短轮训练优于对通用聊天数据进行相同量的训练,对其他医疗提示的纠正与对中毒提示本身的纠正效果大致相当,并且指示纠正编写者模拟一个谨慎、避免伤害的助手,与普通纠正相比没有可测量的额外益处。在我们测试的设置中,纠正有害训练数据比删除它更能减少EM。

英文摘要

Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter's. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.

Comments18 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑