arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MCPEvol-Bench:跨MCP服务器动态演化对大语言模型智能体性能进行基准测试

MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

Huanxi Liu, Kun Hu, Jiaqi Liao, Qiang Wang, Pengfei Qian, YuanZhao Zhai, Dawei Feng, Bo Ding, Huaimin Wang

arXiv 2607.14642首次发表:更新:

发表机构

College of Computer Science and Technology, National University of Defense Technology; State Key Laboratory of Complex & Critical Software Environment; National Key Laboratory of Parallel and Distributed Computing(国防科技大学计算机科学与技术学院; 复杂关键软件环境国家重点实验室; 并行与分布式计算国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对MCP服务器工具接口和功能持续演化致评估有缺陷的问题,引入MCPEvol-Bench基准测试,提出11种变异算子模拟工具演化,对12个先进大语言模型测试,发现前沿模型也难适应,凸显大语言模型驱动工作流程脆弱性,确立评估标准。

AI 中文摘要

随着模型上下文协议(MCP)服务器成为连接大语言模型与外部工具的核心基础设施,现有基准测试利用真实世界的MCP服务器评估大语言模型智能体的工具使用能力。但这些基准测试忽略了MCP服务器中工具接口和功能的持续演化,导致评估有缺陷,无法捕捉智能体在不断变化的工具环境中的适应性。为弥补这一差距,我们引入了MCPEvol-Bench,这是一种用于评估动态工具集演化下大语言模型智能体任务解决能力的新型基准测试。受大规模实证研究启发,我们提出11种变异算子来模拟123个MCP服务器中的实际工具演化。我们在多个版本的MCP服务器上对12个先进大语言模型进行基准测试,发现即使前沿模型也难以适应不断演化的工具。例如,GPT-5.4和Claude-Sonnet-4-6在演化后的MCP服务器中性能分别下降了13.7%和14.4%,同时规划和推理错误大幅增加。这些发现凸显了大语言模型驱动工作流程的脆弱性,确立了MCPEvol-Bench作为评估动态工具环境中智能体适应性的标准。

英文摘要

As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's adaptability in changing tool landscapes. To bridge this gap, we introduce \textbf{MCPEvol-Bench}, a novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution. Inspired by large-scale empirical study, we propose 11 mutation operators to simulate realistic tool evolution within 123 MCP servers. We benchmark 12 state-of-the-art LLMs on multiple versions of MCP servers, revealing that even frontier models struggle to adapt to evolving tools. For instance, GPT-5.4 and Claude-Sonnet-4-6 exhibit performance declines of 13.7\% and 14.4\% in evolved MCP servers, respectively, accompanied by substantial increases in planning and reasoning errors. These findings highlight the vulnerability of LLM-driven workflows, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑