AI 中文总结
本研究引入金丝雀工具诊断LLM智能体的工具选择推理,通过多维度分类法评估8个模型,发现模型能力提升会降低易感性,分类法具能力分层性,且发布了相关框架与数据。
AI 中文摘要
智能体评估显示模型选错了工具,但很少说明原因。我们引入金丝雀工具:植入智能体的模型上下文协议(MCP)工具集中的诊断探测工具,每种工具设计用于探测一种特定的工具选择弱点。六种类型的分类法(语义诱饵、参数陷阱、能力幻象、前提失明、时间诱饵和粒度陷阱)将单一的“选错工具”结果转化为模型如何推理工具的多维度概况。我们评估了8个模型——6个托管模型和2个8B开放权重模型——涵盖三个能力层级,在120个任务上,设置三种金丝雀密度条件和三个随机种子(共8640次运行),外加2880次的微妙性消融实验。任务成功由独立于提供商的评判者评分,并有第二位独立评判者佐证(Cohen's kappa = 0.75)。我们报告三项发现:第一,随着模型能力提升,易感性急剧下降:每任务金丝雀易感性率(CSR)在各模型间相差约36倍,最低为Claude Opus 4.8,最高为Llama 3.1 8B;第二,仅能力层级无法预测安全性:最易受影响的托管模型为中端层级,且在同一提供商内,更便宜的模型可能更安全;第三,分类法具有能力分层性:能力幻象最能有效捕获前沿模型,而其他类型在强模型上基本无效,但会在小型开放模型上触发,因此它们按能力区分,而非仅针对弱模型。弱化每个金丝雀的提示语几乎未改变前沿模型的CSR,证明探测的是推理而非短语识别。易感性也可预测任务失败(Spearman rho = -0.34),而最鲁棒的模型不会因金丝雀压力显著降级。我们发布该框架、金丝雀模式、任务及日志。
英文摘要
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness. A six-type taxonomy (semantic decoys, parameter traps, capability mirages, prerequisite blindness, temporal decoys, and granularity traps) turns a single "wrong tool" outcome into a multi-dimensional profile of how a model reasons about tools. We evaluate eight models -- six hosted and two 8B open-weight -- spanning three capability tiers, on 120 tasks across three canary-density conditions and three seeds (8,640 runs), plus a 2,880-run subtlety ablation. Task success is graded by a provider-independent judge, corroborated by a second independent judge (Cohen's kappa = 0.75). We report three findings. First, susceptibility drops sharply as models get more capable: the per-task canary susceptibility rate (CSR) ranges about 36x across models, lowest for Claude Opus 4.8 and highest for Llama 3.1 8B. Second, capability tier alone does not predict safety: the most susceptible hosted model is mid-tier, and within a provider the cheaper model can be the safer one. Third, the taxonomy is capability-stratified: capability mirages most reliably trap frontier models, while the other types are largely inert on strong models but fire on small open models, so they discriminate by capability rather than being weak. Softening each canary's give-away phrase leaves frontier CSR essentially unchanged, evidence that the probes measure reasoning, not phrase-spotting. Susceptibility also predicts task failure (Spearman rho = -0.34), while the most robust models are not significantly degraded by canary pressure. We release the framework, canary schemas, tasks, and logs.
Comments10 pages, 9 figures, 5 tables