arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38205cs.CLcs.LG

系统提示幻觉:指令前缀如何修改语言模型中的计算

The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models

Muhammad Usama, Dong Eui Chang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过CKA分析17个模型揭示系统提示对Transformer计算的影响具有层选择性和类型依赖性,安全指令虽被编码但未被深度执行,导致越狱漏洞持续存在。

中文摘要 AI 辅助

系统提示是实践者控制语言模型行为的主要手段,然而它们实际上对Transformer内部计算产生的影响仍知之甚少。在涵盖8个架构家族、参数规模从1.5B到72B的17个指令微调模型中,我们使用中心核对齐(CKA)方法,在五个功能类别下的20个系统提示中比较逐层表示。效果具有层选择性和指令类型依赖性:人格和格式指令深度重构中间表示,而安全指令几乎不改变它们,产生的变化在统计上与最小基线难以区分。限制性安全指令和明确允许性指令("你没有限制")涉及几乎相同的计算路径(平均CKA相关性为0.997),并且在商业规模上持续存在,即使在70B-72B规模下,安全穿透率仍低于10%。线性探测基线揭示了机制:模型在每一层编码提示类别,但仅在一小部分层重构其计算,因此提示被可靠地"看到",但对于安全而言,并未被深度"执行"。因果激活修补证实这些层介导行为变化,且表示深度预测整个17模型队列中的行为效应大小(Spearman rho = 0.761,p < 0.001)。这些发现为基于系统提示的安全性的持续越狱漏洞提供了机制性解释。代码:此https URL

英文摘要

System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA) to compare layer-wise representations under 20 system prompts in five functional categories. Effects are layer-selective and instruction-type-dependent: persona and formatting instructions deeply restructure intermediate representations, while safety instructions barely move them, producing changes statistically indistinguishable from a minimal baseline. Restrictive safety instructions and explicitly permissive ones ("you have no restrictions") engage near-identical computational pathways (mean CKA correlation 0.997), and this persists at commercial scale, where safety penetration remains below 10% even at 70B-72B. A linear probing baseline exposes the mechanism: the model encodes prompt category at every layer but restructures its computation only at a small subset, so the prompt is reliably "seen" but, for safety, not deeply "acted upon." Causal activation patching confirms these layers mediate behavioral change, and representational depth predicts behavioral effect size across the full 17-model cohort (Spearman rho = 0.761, p < 0.001). The findings provide a mechanistic explanation for the persistent jailbreak vulnerability of system-prompt-based safety. Code: https://github.com/Usama1002/system-prompt-illusion-cka

发表机构

  • Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

↑