AI 中文总结
该研究通过88.6天对含19099台服务器的MCP注册中心的120次观测,发现按漂移排名重新审计无法充分覆盖描述层面变更,提出内容绑定加定期全目录扫描的控制方案。
AI 中文摘要
对模型上下文协议(MCP)生态系统的安全研究采用了统一设计:每项研究都在单个时间点对注册中心进行审计,却没有一项研究报告这些审计所评判的注册中心描述能保持多久的时效性——这是任何描述层面的发现仍能适用的必要条件(但非充分条件):我们测量的是审计文本的 shelf-life,而非安全发现本身的有效性(第7.1节)。我们在88.6天内重建了官方MCP注册中心的120次观测,覆盖19099台不同服务器,其数量从3510增长至18966。我们的核心结果是一项策略结论:无法通过重新审计漂移最严重的服务器来保持描述层面发现的时效性。在5%的最高重新审计预算下,按过往漂移排名仅能捕获保留窗口内描述发生变化的已观测服务器的约20%(而整体描述漂移的捕获率约为27%),仅能捕获所有描述变更服务器的约10%。这种限制并非源于不可预测性:该排名仍能带来约4倍的提升,而是因为描述层面的变更范围较稀疏——仅8.6%的服务器会重写描述,而描述符的这一比例为24.8%——且约一半的描述变更发生在新加入的服务器上,而历史排名无法覆盖这些服务器,因此相同的提升带来的覆盖范围远小。合适的控制措施是内容绑定:当描述的哈希值变化时立即重新验证,再加上定期的全目录扫描;漂移历史排名最多只是一种部分的、无法覆盖新加入服务器的控制措施。这是描述层面审计员的扫描程序卫生要求,而非运行时信任信号。在至少观测了10个时间间隔的服务器中,四分之三从未变更,最活跃的5%产生了61%的所有变更事件,且同一队列中仅11.9%的描述符在30天内变更;简单的复合模型预测30天内为35.8%(89天内为73%),我们仅将此作为诊断依据,这是一种重尾分布的高估。我们发布了该面板、图表生成器及分析代码。
英文摘要
Security studies of the Model Context Protocol (MCP) ecosystem share a design: each audits a registry at a single point in time. None reports how long the registry descriptions those audits judged stay current - a necessary condition for any description-level finding to still apply, though not a sufficient one: we measure the shelf-life of the audited text, not the validity of a security finding itself (Sec. 7.1). We reconstruct 120 observations of the official MCP registry over 88.6 days, covering 19,099 distinct servers as it grew from 3,510 to 18,966. Our central result is a policy one: you cannot keep description-level findings current by re-auditing the servers that drift most. At a top-5% re-audit budget, ranking by prior drift catches only ~20% of the previously-seen servers whose description changes in a held-out window - versus ~27% for descriptor drift overall - and only ~10% of all description changers. The limit is not unpredictability: the ranking still buys ~4x lift. It is that the description surface is sparse - 8.6% of servers ever rewrite one, against 24.8% for descriptors - and that roughly half of all description changes land on new arrivals a history ranking cannot reach, so the same lift buys far less coverage. The control that fits is content-binding - revalidate the moment a description's hash moves - plus a sized periodic full-catalog sweep; a drift-history ranking is at best a partial, blind-to-new-arrivals control. This is scanner hygiene for a description-level auditor, not a runtime trust signal. Of servers observed across at least ten intervals, three-quarters never change, the most active 5% generate 61% of all change events, and only 11.9% of a cohort's descriptors change within 30 days; naive compounding predicts 35.8% at 30 days (73% at 89), a heavy-tail overestimate we use only as a diagnostic. We release the panel, figure generator, and analysis code.
Comments13 pages, 1 figure. Dataset and code: Zenodo DOI 10.5281/zenodo.21709945 (CC-BY-4.0); this paper version corresponds to dataset v4, DOI 10.5281/zenodo.21798111. v2 corrects five claims found by re-running the paper against its own deposited artifact; each is itemised in the "Changes in v2" section (p.2). The concentration, survival and targeting-coverage findings are unchanged