arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MCP-GRANITE 基准:基于 MCP 的 LLM 智能体粒度接口测试

MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents

Demetris Paschalides, Moysis Symeonides, George Pallis, Marios D. Dikaiakos

arXiv 2609.24161首次发表:更新:

发表机构

University of Cyprus(塞浦路斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MCP智能体工具接口粒度问题,提出MCP-GRANITE基准框架,通过81个场景和8748次试验验证,发现4工具接口最优,任务完成率提升16.4%,并确定粒度为关键设计参数。

AI 中文摘要

随着 LLM 智能体日益通过标准化协议(如 MCP)与外部工具交互,工具接口设计成为一个关键但尚未充分探索的因素。功能如何分解为工具,影响着智能体能否选择正确的工具并构造有效的参数。这一选择在边缘场景中尤为重要,因为资源限制决定了哪些模型可以在本地运行,而扩展规模往往不可行。我们提出了 MCP-GRANITE,一个开源可扩展的基准框架,将工具接口粒度作为基于 MCP 的智能体的受控变量,并在边缘和物联网场景下进行评估。该框架包含 9 个领域的 81 个多步骤场景,实例化为 4 个粒度级别,从细粒度原始工具到单一工具。我们评估了 9 个本地部署的模型(参数规模 268M-20.9B),进行了 8,748 次试验,使用任务完成率、工具选择 F1、参数准确性、延迟和资源使用指标。结果表明,4 工具接口提供了最佳权衡,任务完成率比细粒度原始工具提高 16.4%,比单一整体工具提高 33.6%,同时参数准确性几乎翻倍。模型大小与任务完成率仅弱相关,与延迟强相关,而与参数准确性的关联不太稳健;在最优粒度下,3.2B 模型的表现优于在错配粒度下的 20.9B 模型。这些发现将工具接口粒度确定为基于 MCP 的智能体的关键设计参数。

英文摘要

As LLM agents increasingly interact with external tools through standardized protocols such as MCP, tool-interface design becomes a critical yet underexplored factor. How funψtionality is decomposed into tools affects whether an agent can select the right tool and construct valid arguments. This choice is especially consequential at the edge, where resource constraints limit which models can run locally and scaling up is often not an option. We present MCP-GRANITE, an open-source extensible benchmark framework that treats tool-interface granularity as a controlled variable for MCP-based agents, evaluated under edge and IoT scenarios. It comprises 81 multi-step scenarios across 9 domains, instantiated at 4 granularity levels from fine-grained primitive tools to a single tool. We evaluate 9 locally deployed models (268M-20.9B parameters) across 8,748 trials using task completion, tool selection F1, argument accuracy, latency, and resource-usage metrics. Results show that a 4-tool interface offers the best trade-off, improving task completion by 16.4% over fine-grained primitives and 33.6% over a single monolithic tool, while nearly doubling argument accuracy. Model size is only weakly correlated with task completion and strongly with latency, while its association with argument accuracy is less robust, and a 3.2B model at the optimal granularity outperforms a 20.9B model at a mismatched one. These findings identify tool-interface granularity as a key design parameter for MCP-based agents.

CommentsAuthor copy of paper published at 34th International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication System (MASCOTS2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑