发表机构
William & Mary(威廉与玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建了MCP生态系统的最大工具级跨市场地图MCPacific,通过层级功能分类体系组织超百万工具规范,揭示功能替代普遍但分布不均,且分类展示可显著提升智能体任务完成率。
AI 中文摘要
AI智能体日益依赖通过模型上下文协议(MCP)暴露的工具来完成用户任务。市场上有数十万个MCP服务器被列出,但它们仅按粗略的、特定于市场的服务器类别进行组织。这使得智能体和用户难以识别用于给定操作的工具、找到功能替代方案,并评估这些替代方案之间的差异。我们提出了MCPacific,这是MCP生态系统中最大的工具级、跨市场地图。MCPacific收集了来自17个市场的368,754个MCP服务器列表,对应124,267个唯一服务器,从这些服务器中以七种语言静态提取了1,328,233个工具规范,并将它们组织成一个包含58,915个功能的层级功能分类体系。我们通过迭代的LLM驱动的设计-测试-改进过程构建了该分类体系,并使用校准的嵌入路由将整个语料库映射到其中。我们的研究揭示,MCP的用途远超开发者工具,85%的工具服务于其他领域。功能替代方案普遍存在但分布不均:98.5%的工具至少有一个替代方案,但近四分之一的能仅由一个工具支持。功能上可比的工具在安全警报、代码复杂度和项目维护方面也存在差异,在41%的可比工具对中,复杂度差异超过2.5倍。最后,通过分类体系而非平面列表展示候选工具,在所有四个评估模型上提高了任务完成率,在拥挤的候选集上,Pass@0.75的提升高达12个百分点。
英文摘要
AI agents increasingly rely on tools exposed through the Model Context Protocol (MCP) to complete user tasks. Hundreds of thousands of MCP servers are listed across marketplaces, yet they are organized only by coarse, marketplace-specific server categories. This makes it difficult for agents and users to identify tools for a given operation, find functional alternatives, and assess how those alternatives differ. We present MCPacific, the largest tool-level, cross-marketplace map of the MCP ecosystem. MCPacific collects 368,754 MCP server listings corresponding to 124,267 unique servers across 17 marketplaces, statically extracts 1,328,233 tool specifications from these servers in seven languages, and organizes them into a hierarchical functional taxonomy of 58,915 capabilities. We construct the taxonomy through an iterative LLM-driven design-test-refine process and map the full corpus to it using calibrated embedding routing. Our study reveals that MCP extends well beyond developer tooling, with 85% of tools serving other domains. Functional alternatives are widespread but unevenly distributed: 98.5% of tools have at least one alternative, yet nearly a quarter of capabilities are supported by only one tool. Functionally comparable tools also differ in security alerts, code complexity, and project maintenance, with complexity differing by more than 2.5x in 41% of comparable tool pairs. Finally, presenting candidate tools through the taxonomy rather than a flat list improves task completion rate across all four evaluated models, with gains of up to 12 percentage points in Pass@0.75 for crowded candidate sets.