发表机构
National Engineering Research Center for Software Engineering; Peking University; City University of Hong Kong; Tencent Technology(国家软件工程研究中心; 北京大学; 香港城市大学; 腾讯科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现延续框架是上下文学习涌现失配的关键调节因素,演示框架会使Gemini模型的涌现失配提升30至32个百分点,且不同模型对此的响应存在差异。
AI 中文摘要
上下文学习(ICL)可引发涌现失配(EM),即窄范围的失配示例会改变对不相关问题的回答。然而现有提示将有害文本暴露与邀请延续助手行为混为一谈。我们固定有害答案,仅改变其作为演示、证据、助手历史或工具输出的呈现形式。在10个独立采样的上下文里,演示框架使易受影响的Gemini模型上的广义EM提升了30至32个百分点;该差距在排除领域、语义聚类、未见过的问题及4种提示模板后仍存在。格式和长度匹配的对照实验表明,有害内容是必要但非充分条件。角色×延续的析因实验进一步揭示了模型依赖的来源效应:Gemini会遵循助手和工具历史,而Grok大多抵制工具框架的延续;其他几个前沿及开放权重模型则无此差距。盲法人工审核确认了所有主要对比,且模型评判低估了活跃条件下的失败情况。因此,延续框架是ICL-EM的强模型依赖调节因素,而非有害上下文的普遍后果。
英文摘要
In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts, however, conflate harmful-text exposure with an invitation to continue assistant behavior. We hold harmful answers fixed while varying their delivery as demonstrations, evidence, assistant history, or tool output. Across ten independently sampled contexts, demonstration framing raises broad EM by $30$--$32$ percentage points on a susceptible Gemini model; the gap survives domain exclusion, semantic clustering, unseen questions, and four prompt templates. Format and length-matched controls show that harmful content is necessary but insufficient. A role times continuation factorial further reveals model-dependent provenance effects: Gemini follows both assistant and tool histories, whereas Grok largely resists tool-framed continuation. Several other frontier and open-weight models show no gap. Blinded human audits confirm every main contrast and show that the model judge underestimates active-condition failures. Thus continuation framing is a strong, model-dependent moderator of ICL-EM, not a universal consequence of harmful context.