发表机构
Iowa State University; Meta(爱荷华州立大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过机制可解释性分析角色扮演越狱中模型服从有害请求的机制,发现安全继电器衰减等规律,明确了安全措施需维持有害识别到拒绝的关联。
AI 中文摘要
大型语言模型经训练后会遵循指令并拒绝有害请求,越狱技术会利用这种平衡诱导模型生成通常会拒绝的内容,其中角色扮演越狱尤其值得关注:有害请求仍可见于由角色设定、场景和任务构成的角色扮演包装中,但模型可能仍会服从。本研究采用机制可解释性方法,探究该上下文如何逆转拒绝以及哪些元素促成了这种逆转。在两个基准、三个模型家族和四个原创包装上,我们对比了带有和不带该包装的匹配有害与良性请求,追踪从请求到最终提示状态的隐藏状态对比,通过受控反事实隔离包装操作,在留存评估请求中干预其激活方向,并通过几何方法分解有效方向。分析得出三项发现:(1)成功攻击在请求阶段保留了测得的有害-良性区分,但在答案开始处与拒绝相关的表达减弱,该模式被称为安全继电器衰减;(2)围绕请求构建完整角色扮演并在场景中构建框架具有因果作用:移除相关激活变化可恢复拒绝;(3)这些效应大多共享内部结构,且大多数修复可通过与模型无角色扮演时对有害请求的常规拒绝对齐的组件实现,场景框架保留了较小的、依赖模型的组件。综上,这些发现解释了角色扮演为何能在保留有害证据的情况下产生服从行为,并为未来安全措施确定了具体目标:维持从有害识别到拒绝的关联。
英文摘要
Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply. We use mechanistic interpretability to determine how this context reverses refusal and which elements contribute to the reversal. Across two benchmarks, three model families, and four authored wrappers, we compare matched harmful and benign requests with and without this wrapper. We trace hidden-state contrasts from the request to the final prompt state, isolate wrapper operations through controlled counterfactuals, intervene on their activation directions in held-out evaluation requests, and decompose effective directions geometrically. Our analysis yields three findings. (1) Successful attacks retain the measured harmful-versus-benign distinction at the request, while its refusal-associated expression weakens where the answer begins, a pattern we call safety-relay attenuation. (2) Constructing the complete roleplay around the request and framing it within the scenario contribute causally: removing the associated activation changes restores refusal. (3) These effects largely share internal structure, and most repair is reproduced by components aligned with the model's ordinary refusal of harmful requests without roleplay; scenario framing retains a smaller, model-dependent component. Together, these findings explain how roleplay can produce compliance despite retained evidence of harm and identify a concrete target for future safeguards: maintaining the connection from harm recognition to refusal.
CommentsPreprint