发表机构
Zenity(Zenity)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究LLM智能体激活探针监控中,在用户回合后附加分类提示后缀能否提升野外恶意输入检测泛化能力,发现分类格式而非具体标准带来增益,且效果依赖模型与读取方式,为部署提供廉价改进方案。
AI 中文摘要
LLM智能体越来越依赖激活探针作为运行时监控器,用于检测提示注入、越狱和不安全请求,在智能体对有害输入采取行动之前读取模型自身的隐藏状态以捕获有害输入。一种廉价且日益常见的做法,借鉴自LLM-as-judge提示技术,是在用户回合后附加一条简短的分类指令,并在该点读取探针以使其更敏锐:该指令要求模型将传入请求表示为某个类别,集中探针必须分离的信号,且服务成本可忽略不计。但是,该后缀的措辞是否重要,其益处是否在野外场景中——即面对探针在训练中从未见过的攻击类型,即部署监控器所面临的场景——依然成立?我们通过一系列受控的用户后后缀阶梯,在严格的留一数据集外(LODO)评估下,跨13个安全基准(越狱、注入和良性聊天)和三个开放权重模型家族(Llama-3.1-8B、Qwen3.5-9B、Gemma-4-12B)进行了测试。在单位置探针上,分类后缀相比无后缀持续改善了分布外检测(最高约4个AUC点);然而,哪个后缀起作用很重要:提示模型对输入进行分类,即使使用无内容标签,也可靠地胜出;而离题或仅注意的后缀帮助甚微。增益来自分类格式,而非命名标准:无内容后缀与真实的恶意/良性后缀相匹配,标准仅在严格阈值下增加精确度。这不是单位置读取的伪影:益处延续到生产中使用的多位置池化探针(注意力、多最大值、MLP),尽管那里表现最佳的后缀取决于读取方式。通过KV缓存分叉提供服务,它是任何激活探针监控器的廉价即插即用方案,但并非自动获胜:哪个后缀有帮助以及帮助多少,取决于模型和读取方式。
英文摘要
LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to append a short classification instruction after the user's turn and read the probe at that point, to sharpen it: the instruction asks the model to represent the incoming request as a class, concentrating the signal the probe must separate, at negligible serving cost. But does the wording of that suffix matter, and does its benefit hold in the wild, on attack types the probe never saw in training, the regime a deployed monitor faces? We test this with a controlled ladder of post-user suffixes under strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B). On a single-position probe, a classification suffix consistently improves out-of-distribution detection over no suffix (up to ~4 AUC points); yet which suffix matters: prompting the model to classify the input, even into content-free labels, reliably wins; an off-topic or merely-attentive suffix helps little. The gain comes from the classification format, not the named criterion: a content-free suffix matches the real malicious/benign one, with the criterion adding precision only at strict thresholds. This is not an artifact of the single-position read: the benefit carries to the multi-position pooling probes used in production (attention, multi-max, MLP), though the best-performing suffix there is readout-dependent. Served through a KV-cache fork, it is a cheap drop-in for any activation-probe monitor, though not an automatic win: which suffix helps, and by how much, depends on the model and the readout.
Comments40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: Agents in the Wild: Safety, Security, and Beyond