发表机构
Old Dominion University(奥尔德多米宁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
COPEX是一个受控基准测试,隔离MCP系统中模型作为客户端的对抗性鲁棒性,涵盖25种攻击类型,平均攻击成功率为64.4%,并验证输入与上下文扫描可显著降低攻击成功率。
AI 中文摘要
大型语言模型越来越多地在模型上下文协议(MCP)系统中作为工具使用的中介,在此类系统中,对抗性影响可能通过用户指令、工具模式、工具输出或协议消息进入。现有基准测试通常评估已部署的智能体,将模型易感性与护栏、编排和通用任务能力混为一谈。我们提出了COPEX(上下文提供者利用),一个受控基准测试,通过固定周围的智能体栈并仅改变工具选择模型,将模型隔离为MCP客户端。COPEX涵盖25种攻击类型,实例化为跨四个入口表面(模型/智能体、客户端、服务器/工具和传输)的125个场景。在九个模型和3,375次试验中,平均攻击成功率为64.4%,各表面平均成功率范围从58.3%到71.4%。一些客户端和传输级别的攻击部分在模型的观察或控制之外成功,将系统暴露与模型易感性区分开来。结合输入和上下文扫描,在八种攻击的防御子集上,相对于未防御设置,平均攻击成功率降低了49.6%。该基准测试可在https://this-url上获取。
英文摘要
Large language models increasingly mediate tool use in Model Context Protocol (MCP) systems, where adversarial influence may enter through user instructions, tool schemas, tool outputs, or protocol messages. Existing benchmarks often evaluate deployed agents, conflating model susceptibility with guardrails, orchestration, and general task capability. We introduce COPEX (COntext Provider EXploitation), a controlled benchmark that isolates the model as an MCP client by fixing the surrounding agent stack and varying only the tool-selecting model. COPEX covers 25 attack types instantiated as 125 scenarios across four entry surfaces: model/agent, client, server/tool, and transport. Across nine models and 3,375 trials, the mean attack success rate is 64.4%, with surface-level means ranging from 58.3% to 71.4%. Some client- and transport-level attacks succeed partly outside the model's observation or control, separating system exposure from model susceptibility. Combined input and context scanning reduces mean attack success by 49.6% on an eight-attack defense subset relative to the undefended setting. The benchmark is available at https://github.com/inspire-center/copex.
CommentsAccepted as a poster at the NeurIPS 2026 Workshop on Agents in the Wild