AI 中文总结
研究对1723个GitHub上的MCP应用展开大规模研究,先从样本得出分类,再用大语言模型辅助流程刻画服务器集成。结果表明该生态系统部分实践趋同,如多数用文件配置服务器、用官方SDK通信,但配置文件命名无约定,人工监督差异大。
AI 中文摘要
模型上下文协议(MCP)规范了大语言模型应用与外部工具的通信方式,但应用端未作明确规定。与通过包管理器解决的传统依赖不同,集成MCP服务器的开发者在配置、通信或人工监督方面没有约定。现有工作多关注服务器而非使用它们的应用,该生态系统研究不足。我们对从GitHub挖掘的1723个MCP应用进行了大规模研究。先从代表性样本中得出MCP应用分类,再用大语言模型辅助流程应用于整个数据集,刻画服务器集成情况。结果显示,该生态系统在一些实践上已趋同,但在另一些方面并非如此。多数MCP应用通过文件配置服务器、使用官方软件开发工具包与服务器通信,但配置文件尚无命名约定。人工监督差异最大,日志记录和启用/禁用控制很常见,但只有37.2%在阻止审批步骤后控制工具执行,多数MCP应用中,大语言模型能无条件调用任何启用的工具。
英文摘要
The Model Context Protocol (MCP) standardizes how large language model applications communicate with external tools, but leaves the application side unspecified: unlike traditional dependencies resolved through package managers, developers integrating MCP servers face no conventions for configuration, communication, or human oversight. This ecosystem is also under-researched, with existing work focused on servers rather than the applications consuming them. We conduct a large-scale study of 1,723 MCPApps mined from GitHub. We first derive MCPAppTax from a representative sample, then use an LLM-assisted pipeline to apply it across the full dataset, characterizing server integration across configuration, SDK use, and human-in-the-loop mechanisms. Our results show that the ecosystem has converged on some practices but not others: most MCPApps configure servers using files (85.2%) and use an official SDK (81.1%) to communicate with servers, yet no naming convention has emerged for configuration files. Human oversight diverges most, logging (90.8%) and enable/disable controls (77.2%) are common, but only 37.2% gate tool execution behind a blocking approval step, leaving the LLM able to invoke any enabled tool unconditionally in most MCPApps.