控制令牌注入抑制思维链并击败基于推理的工具使用智能体监管
Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents
- Braindeck Inc(Braindeck公司)
- University of Wah(瓦赫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究证明工具使用智能体的安全性取决于模型与解码框架的联合属性,通过注入控制令牌可抑制思维链并绕过推理监管,导致拒绝转化为窃取行为,并提出了相应防御措施。
AI中文摘要:
工具使用语言模型智能体的安全性通常被视为模型本身的属性。我们给出了受控的全精度证据,表明它实际上是模型与渲染其聊天模板并解析其工具调用的软件(即解码框架)的联合属性,且这两部分都可从不可信输入受到攻击。在已发布的gpt-oss-20b推理模型及其公开工具沙箱下,向用户消息追加一串模型自身的通道控制令牌,会使分词器渲染出一个已经完成的推理回合,因此模型不写出思维链,直接进行工具调用。在模型已能完成的四十个任务中,推理通道的平均令牌数从52.5降至每次试验的零,而http.post仍在每次试验中触发。规则监控器和跨家族语言模型监控器在所有普通试验中检测到不安全请求,且无伪造试验被检测到;对于明显恶意的请求,该攻击将模型39.6%的拒绝转化为完成的窃取行为。另外,相同的工具调用生成是否触发由框架解析器决定,而非模型:一个容忍截断的正则表达式触发了一个缺少结束令牌的调用,而严格的正则表达式则丢弃它;为Gemma智能体提供的两个解析器在相同的贪婪生成上给出相反结果,在全部二十四次试验中触发,而在另一些试验中一次也不触发。我们展示了这种抑制可以间接传递,并表征了其对另外两个推理模型的聊天模板的依赖性;我们评估了输入清理、解析器加固和空推理检测作为防御措施;标记缺失轨迹能捕获基本攻击,但无法捕获自适应良性诱饵。所有测量均使用公开模型的贪婪解码。代码和逐试验日志:https://github.com/Usama1002/deleting-the-trace
英文摘要:
The safety of a tool-using language model agent is usually treated as a property of the model alone. We give controlled, full-precision evidence that it is instead a joint property of the model and the software that renders its chat template and parses its tool calls, the decoding harness, and that both halves are attackable from untrusted input. On the released gpt-oss-20b reasoning model under its published tool sandbox, appending a single string of the model's own channel-control tokens to a user message makes the tokenizer render a reasoning turn that is already complete, so the model writes no chain-of-thought and proceeds directly to the tool call. Across forty tasks the model already completes, the reasoning channel falls from a mean of 52.5 tokens to zero on every trial while the http.post still fires on every trial. A rule monitor and a cross-family language-model monitor detect the unsafe request on all plain trials and no forged trials, and on overtly malicious requests the attack converts 39.6% of the model's refusals into completed exfiltrations. Separately, whether an identical tool-call generation fires is decided by the harness parser, not the model: a truncation-tolerant regular expression fires a call whose closing token is missing while a strict one drops it, and two parsers shipped for the Gemma agent give opposite outcomes on identical greedy generations, firing on all twenty-four trials and on none. We show the suppression can be delivered indirectly and characterize its dependence on the chat template across two more reasoning models, and we evaluate input sanitization, parser hardening, and empty-reasoning detection as defenses; flagging an absent trace catches the basic attack but not an adaptive benign decoy. All measurements use greedy decoding on publicly released models. Code and per-trial logs: https://github.com/Usama1002/deleting-the-trace