发表机构
University of Southern California (USC)(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对作为智能体护栏的类型化决策模型提出选项通道攻击,发现其在提示注入等任务中准确率低,易被6行无关文本或误导性选项名称突破,现有防御失效,确定性规则可替代模型实现100%准确率。
AI 中文摘要
类型化决策模型读取一段文本,返回调用者定义的选项上的概率分布,每个选项带有简短的书面定义,不生成文本。近期研究将这些模型置于智能体系统中作为护栏:该组件读取提议的工具调用或传入消息,并决定是否允许其执行。我们评估了7个开放权重模型在该角色下的表现,并分别报告两种错误方向:开放失败错误允许被禁止的操作,属于漏洞;封闭失败错误阻止被允许的操作,仅造成成本。在提示注入、越狱和有毒内容筛选任务中,这些模型在允许或阻止决策上的准确率在36%至72%之间,而随机水平为50%。仅某一方向的低错误率反映了模型的默认答案:一个模型几乎允许所有内容,另一个几乎阻止所有内容。在一组合成的智能体工具调用中,6行与策略无关的服务器日志文本,使原本能正确判断的策略的闸门开放失败率从0%升至63%。在4个将标签置于输入中的模型上,仅将允许性选项的名称改为误导性名称,而其定义和待判断文本保持不变,就使该开放失败率升至93%至100%。我们测试的所有防御均被击败,要么是攻击者针对其机制,要么是攻击者控制的文本。提高置信度最低的决策也无济于事:被攻击反转的决策与被替换的决策置信度相同。将每个策略字段解析为类型化值确实消除了一种攻击,但也使模型变得不必要:对这些值的确定性规则在所有6个策略上达到100%准确率。这些模型可以减少到达审核者的案例数量,但根据现有证据,它们不应成为决策组件。代码可在该https URL获取。
英文摘要
A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text. Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it. We evaluate seven open-weight models in that role and report the two error directions separately: a fail-open error allows a prohibited action and is a vulnerability; a fail-closed error blocks a permitted one and is only a cost. On prompt-injection, jailbreak and toxic-content screening, accuracy at the allow-or-block decision ranges from 36% to 72% against a chance level of 50%. A low error rate in one direction only reflects which answer a model defaults to: one allows nearly everything, another blocks nearly everything. On a synthetic suite of agent tool calls, six lines of server log text that say nothing about the policy raise a gate's fail-open rate from 0% to 63% on a policy it otherwise decides correctly. Giving the permissive option a misleading name, with its definition and the judged text untouched, raises that rate to between 93% and 100% on the four models that place the label in their input. Every defense we tested is defeated, either by an attacker who targets its mechanism or by attacker-controlled text. Escalating the least confident decisions does not help either: a decision an attack has reversed is no less confident than the one it replaced. Parsing each policy field into a typed value does eliminate one attack, but it also makes the model unnecessary: a deterministic rule over those values reaches 100% accuracy on all six policies. These models can reduce how many cases reach a reviewer, but on this evidence they should not be the component that decides. Code is available at https://github.com/ArminAzizi98/option-channel-attack.