arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考MCP安全性:对运行时MCP服务器和安全扫描器可靠性的大规模研究

Rethinking MCP Security: A Large-Scale Study of Runtime MCP Servers and Security Scanner Reliability

Pei Chen, Baichao An, Mengying Wu, Binwang Wan, Geng Hong, Jinsong Chen, Xudong Pan, Jiarun Dai, Min Yang

arXiv 2607.11086首次发表:更新:

AI 中文总结

研究重新审视MCP安全性衡量方式,通过多智能体框架构建MCPZoo进行大规模测量。发现现有扫描器报告的风险信号不可靠,人工验证抽样警报真阳性率低且扫描器输出不一致。实现对MCP服务器安全性的大规模、可重复测量并揭示扫描局限,还发布公共查询接口。

AI 中文摘要

模型上下文协议(MCP)已迅速成为基于大语言模型的代理与外部工具和服务交互的标准接口。随着MCP服务器越来越多地被委托进行安全敏感操作,了解其实际风险变得至关重要。由于缺乏大规模运行时MCP服务器,这种理解很大程度上依赖于应用于少数案例的安全扫描器,但其评估的可靠性仍不清楚。本研究重新审视了MCP安全性的衡量方式。我们展示了MCPZoo,这是迄今为止用于动态分析的最大的MCP服务器集合。MCPZoo通过多智能体框架构建,将野外静态存储库转换为动态服务。该框架通过将环境推理与反馈驱动的优化相结合,模拟人类专家构建、诊断和迭代修复部署及运行时缺陷的方式。为确保运行时的实际交互性,服务器通过真实协议交互进行验证。结果,MCPZoo包含64,611个独特的MCP服务器(总共113,927个),其中超过37,288个支持动态分析。利用MCPZoo,我们首次对MCP服务器及其分析扫描器进行了生态系统规模的测量。虽然现有扫描器报告称96.89%的服务器存在风险,但我们发现这些信号不可靠。特别是,人工验证表明,抽样警报中不到50%是真阳性,并且扫描器输出在不同扫描器之间表现出明显的不一致。总体而言,MCPZoo实现了对MCP服务器安全性的大规模、可重复测量,并揭示了当前扫描实践的局限性。我们还发布了一个公共查询接口,以支持对MCP服务器的实际风险评估。

英文摘要

The Model Context Protocol (MCP) has rapidly established itself as a standard interface for enabling LLM-based agents to interact with external tools and services. As MCP servers are increasingly entrusted with security-sensitive operations, understanding their real-world risks has become critical. In practice, due to the absence of large-scale runtime MCP servers, such understanding largely relies on security scanners applied to a small number of cases, yet the reliability of these assessments remains unclear. In this study, we revisit how MCP security is measured. We present MCPZoo, the largest collection of MCP servers for dynamic analysis to date. MCPZoo is constructed through a multi-agent framework for transforming in-the-wild static repositories into dynamic services. The framework emulates how human experts build, diagnose, and iteratively repair deployment and runtime defects by combining environment inference with feedback-driven refinement. To ensure practical interactivity at runtime, the servers are validated via real protocol interactions. As a result, MCPZoo contains 64,611 unique MCP servers (113,927 in total), with more than 37,288 supporting dynamic analysis. Leveraging MCPZoo, we conduct the first ecosystem-scale measurement of MCP servers and the scanners that analyze them. While existing scanners report that 96.89% of servers are risky, we find that these signals are unreliable. In particular, manual validation shows that less than 50% of sampled alerts are true positives, and scanner outputs exhibit clear inconsistency across scanners. Overall, MCPZoo enables large-scale, reproducible measurement of MCP server security and exposes limitations of current scanning practices. We further release a public query interface to support practical risk assessment of MCP servers.

Comments18 pages, 11 figures, and 10 tables. This article substantially extends the preliminary 3-page MCPZoo dataset release arXiv:2512.15144. Includes appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑