发表机构
ITMO University(俄罗斯圣彼得堡国立信息技术机械与光学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLM智能体在MCP服务器上的评估基准问题,提出DynamicMCPBench框架,能生成目标、追踪轨迹并按效果评分。大规模测试发现智能体处理长多步任务能力不足,且该框架可重复运行,人工验证其自动评分可靠。
AI 中文摘要
大语言模型(LLM)智能体越来越多地部署在模型上下文协议(MCP)服务器上,但用于评估它们的基准测试是对最终答案或固定的“地面真值”工具列表进行评分,一旦基础数据是实时且有状态的,这两者都很脆弱。我们提出了DynamicMCPBench,这是一个可重复使用的框架而非固定数据集。从业者可以在自己的MCP服务器上运行它来测试模型在自身任务上的表现,或者让它自动收集服务器以测量模型解决智能体任务的一般能力。给定服务器和任何一组模型,它会生成现实目标,实时追踪每个目标以记录成功轨迹,将该轨迹提炼为与路径无关的效果检查点,并根据智能体是否重现这些效果而非最终答案对其进行评分。为展示该框架所揭示的内容,我们大规模运行它:在121台服务器上对24个模型进行测试,涵盖750个任务,均匀分布在15个任务类别(每个类别50个)中,每个类别针对生成问题的不同工具使用挑战。每个任务通过pass^3评分:只有所有三次独立尝试都成功才算解决。即使是最强的智能体也只能解决约一半任务,31%的任务没有模型能解决,随着所需工具链变长,准确率会下降(从最短链的39%降至最长链的13%)。一项人工验证研究证实自动评分是可靠的(机会校正一致性为0.76)。DynamicMCPBench因此将基准测试构建转变为从业者可以在自己的服务器和模型上重新运行的操作,同时揭示了当前智能体在处理长的、多步骤智能体任务方面持续存在的能力不足。
英文摘要
Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful. We present DynamicMCPBench, a reusable framework rather than a fixed dataset. A practitioner can run it on their own MCP servers to test models on their own tasks, or let it collect servers automatically to measure a model's general ability to solve agentic tasks. Given the servers and any set of models, it generates realistic goals, pursues each one live to record a successful trajectory, distills that trajectory into path-agnostic effect checkpoints, and scores an agent on whether it reproduces those effects, never on the final answer. To show what the framework reveals, we run it at scale: 24 models over 121 servers and 750 tasks spread evenly over 15 task categories (50 each), where each category targets a distinct tool-use challenge of the generated questions. Each task is scored by pass^3: it counts as solved only if all three independent attempts succeed. Even the strongest agents solve only about half of the tasks, 31% of tasks are solved by no model at all, and accuracy collapses as the required tool chain grows longer (from 39% on the shortest chains to 13% on the longest). A human validation study confirms the automatic scoring is reliable (chance-corrected agreement of 0.76). DynamicMCPBench thus turns benchmark construction into something practitioners can rerun on their own servers and models, while exposing a consistent inability of current agents to handle long, multi-step agentic tasks.
CommentsAccepted to EMNLP 2026