arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从 MCP 注册表中随机抽取的内容,以及工具使用基准测试所包含的内容

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

Haseeb Mohammed Afsar

arXiv 2609.10962首次发表:更新:

AI 中文总结

本研究通过概率抽样揭示MCP注册表真实状态:仅48.8%服务器可启动,硬一致性高但安全注释缺失严重,并发现BFCL基准存在大量重复,而真实工具重复率极低。

AI 中文摘要

对模型上下文协议(MCP)服务器生态系统的研究,其样本抽取方式会悄然偏向于选择那些可正常工作的服务器:参考集、流行度列表、手工策划的框架,或修复服务器直至其启动的流水线。我们报告了未经修复的概率样本实际包含的内容。从包含24,135个服务器的注册表普查中,我们抽取了400个带有已发布种子的npm/stdio服务器,并通过网络对每个服务器进行探测。仅48.8%的服务器完成了初始化握手,而使用相同工具测量的手工策划框架的完成率为66.7%,主要失败原因并非缺少凭据(13.3%),而是服务器根本无法启动(37.5%)。在195个实际运行的服务器中,硬一致性是全面的:在2,766个广告工具中,JSON Schema致命违规为零。可选的 safety 注释才是真正的差异来源,随机抽取中工具级遗漏率为58.8%,而策划框架中为41.5%,因此策划也美化了这一数字。然后,我们使用一种保持不变的方法,将这些服务器广告的工具描述与两个工具使用基准语料库进行比较。真实MCP工具在余弦相似度0.70时显示出2.8%的近似重复,且所有这些重复都发生在单个服务器内部:在测试的所有阈值下,跨作者的近似重复率为0.0%。BFCL v4显示16.7%的近似重复,其中16.4个百分点发生在独立呈现的任务之间。UltraTool显示0.3%,比真实工具更干净,因此这是BFCL的特性,而非合成语料库的类别特性。另外,原始BFCL行中68.8%和原始UltraTool行中85.6%是名称加描述的精确重复,而真实MCP中为0.4%,因此,在没有全局去重的情况下,对这些发布版本计算的任何统计量衡量的是重复而非工具。所有数据均可从已发布的脚本和已发布的种子重新生成。

英文摘要

Studies of the Model Context Protocol (MCP) server ecosystem draw their samples in ways that quietly select for servers that work: reference sets, popularity lists, hand-curated frames, or pipelines that repair a server until it starts. We report what an unrepaired probability sample actually contains. From a 24,135-server registry census we draw 400 npm/stdio servers with a published seed and probe each one over the wire. Only 48.8% complete an initialize handshake, against 66.7% for a hand-curated frame measured with the same instrument, and the dominant failure is not missing credentials (13.3%) but servers that never start at all (37.5%). Among the 195 that do run, hard conformance is total: zero fatal JSON Schema violations across 2,766 advertised tools. Optional safety annotations are the real variance, and the tool-level omission rate on a random draw is 58.8% against 41.5% on the curated frame, so curation flatters this figure too. We then compare the tool descriptions these servers advertise against two tool-use benchmark corpora using one method held constant. Real MCP tools show 2.8% near-duplication at cosine 0.70, and all of it lies within single servers: cross-author near-duplication is 0.0% at every threshold tested. BFCL v4 shows 16.7%, of which 16.4 points lie between independently presented tasks. UltraTool shows 0.3%, cleaner than real tools, so this is a property of BFCL and not of synthetic corpora as a class. Separately, 68.8% of raw BFCL rows and 85.6% of raw UltraTool rows are exact name-plus-description repeats, against 0.4% for real MCP, so any statistic computed over these releases without global deduplication measures repetition rather than tools. All figures regenerate from released scripts and a published seed.

Comments10 pages, 3 figures. Seeded, re-runnable pipeline and per-server outcomes: https://github.com/itguruhaseeb/mcp-probe ; archived at doi:10.5281/zenodo.21347997

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑