发表机构
KAIST; INEEJI(韩国科学技术院; 因尼吉)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型中不完整提示越狱现象,通过实证表征其何时及如何引发有害延续,分析相关吸引子类型,指出参数调整训练模型拒绝不完整有害提示不足,识别终止和延续神经元,强调神经元级干预对防御的潜力。
AI 中文摘要
大语言模型(LLMs)越来越多地作为具有防止有害请求保护措施的开放权重模型发布。然而,句子完成仍然容易受到不完整有害提示的影响。在这项工作中,我们将这种现象形式化为不完整提示越狱(IPJ),并对不完整提示何时以及如何引发有害延续进行了系统的实证表征。我们分析了与不完整句子延续相关的不同吸引子类型,并表明大语言模型会系统地延迟拒绝直到句子结束。我们进一步证明,通过参数调整训练模型拒绝不完整有害提示是不够的,无法在内容领域和吸引子类型之间进行泛化。为了实现细粒度控制,我们识别了两个功能神经元:终止和延续神经元。通过阐明它们在句子完成中的作用,我们强调了神经元级干预对于更精确和强大的IPJ防御的潜力。
英文摘要
Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize this phenomenon as incomplete prompt jailbreaks (IPJ) and provide a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations. We analyze diverse attractor types associated with incomplete sentence continuation and show that LLMs systematically delay refusal until sentence termination. We further demonstrate that training models to refuse incomplete harmful prompts via parameter tuning is insufficient, failing to generalize across both content domains and attractor types. To enable fine-grained control, we identify two functional neurons: termination and continuation neurons. By clarifying their roles in sentence completion, we highlight the potential of neuron-level interventions for more precise and robust IPJ defenses.
CommentsAccepted to ACL 2026 Findings. 15 pages (9 pages for main body), 13 figures