arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对齐数据可通过上下文混淆引发不对齐

Aligned Data Can Induce Misalignment via Context Confusion

Yavuz Bakman, Duygu Nur Yaldiz, Baris Askin, Swastik Roy, Morteza Ziyadi, Salman Avestimehr, Sai Praneeth Karimireddy

arXiv 2609.38379首次发表:更新:

发表机构

University of Southern California; Carnegie Mellon University; Amazon; Microsoft(南加州大学; 卡内基梅隆大学; 亚马逊; 微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现对齐训练数据可能通过上下文混淆引发其他上下文中的不对齐行为,提出定向数据与上下文示例可缓解,并揭示其机制,强调训练后评估的重要性。

AI 中文摘要

大型语言模型(LLMs)经常因各种用例而更新,其中过滤掉不对齐的训练样本是防止更新后不对齐的常见做法。然而,对齐本质上是依赖于上下文的:在一个上下文中对齐的建议在另一个上下文中可能不合适。例如,对于问题“研究人员应该如何处理研究数据?”,建议研究人员保留数据以确保可复现性是对齐的。相反,对于“移动应用开发者应该如何处理用户的敏感数据?”,建议保存数据从隐私角度来看可能是不合适的。基于这一观察,我们识别出一种训练后现象,即对齐训练在其他上下文中引发不对齐行为。我们将这种现象称为**上下文混淆**。我们在三个领域展示了上下文混淆:(1)性别平等,(2)隐私,以及(3)人身安全。我们进一步表明,上下文混淆导致狭窄不对齐,与涌现不对齐形成对比,并且通过注入一般对齐数据不能有效减少,但通过包含针对不对齐领域的定向对齐数据或在推理期间提供上下文学习示例可以大幅减少。最后,我们提供了*上下文混淆*的机制解释。我们观察到,来自不同领域的查询在微调期间可能经历相似的表征转变。因此,来自不同领域的查询可能激活在微调期间学到的相同行为特征,这导致行为转移到其不对齐的上下文中。基于我们的发现,我们认为仅通过检查训练数据很难预测模型训练后的对齐状态,这凸显了全面的训练后对齐评估的重要性。

英文摘要

Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher preserve the data for reproducibility is aligned. In contrast, recommending data saving in response to "What should a mobile-app developer do with users' sensitive data?" may be inappropriate from a privacy perspective. Starting from this observation, we identify a post-training phenomenon where aligned training induces misaligned behavior in other contexts. We call this phenomenon **context confusion**. We demonstrate context confusion across three domains: (1) Gender Equality, (2) Privacy, and (3) Physical Safety. We further show that context confusion causes narrow misalignment, in contrast to emergent misalignment, and is not effectively reduced by injecting general alignment data, but can be substantially reduced by including targeted alignment data for the misaligned domain or providing in-context learning examples during inference. Lastly, we provide a mechanistic explanation of *context confusion*. We observe that queries from different domains can undergo similar representational shifts during the fine-tuning. Consequently, a query from a different domain may activate the same behavioral feature learned during fine-tuning, which causes the behavior to transfer to a context where it is misaligned. Based on our findings, we argue that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑