arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cartograph:面向AI代理的基于操作者认证检索的联邦工具发现

Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents

Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong, Jerry John Kponyo

arXiv 2609.30293首次发表:更新:

发表机构

Kwame Nkrumah University of Science and Technology; Ghana Communication Technology University(夸梅·恩克鲁玛科技大学; 加纳通信技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Cartograph提出联邦MCP代理,通过操作者认证卡片、三层混淆簇分析和两阶段检索,将工具发现从O(n)降至O(k),在374工具部署中仅暴露3个代理工具,R@5达0.816,显著减少令牌开销并记录描述来源。

AI 中文摘要

模型上下文协议(MCP)使AI代理能够发现并调用工具,但随着连接目录的增长,加载每个定义变得代价高昂。我们提出Cartograph,一个联邦MCP代理,它将代理可见的工具发现从$O(n)$目录遍历转变为$O(k)$渐进式披露。Cartograph结合了三种机制:(1)操作者认证的能力卡片,即在部署操作者控制下生成的Ed25519签名描述,而非排名后的发布者文案;(2)Rift,一种三层混淆簇分析,包括密度聚类、查询边际分析和令牌诊断;(3)两阶段检索,先对服务器排序再对工具排序。在一个包含22台服务器、374个工具的部署中,Cartograph仅暴露三个代理工具而非374个定义。一个由作者构建的49查询基准实现了R@5为0.816,而Jaccard关键词基线为0.592,同时测量的前五发现交换在所述全目录核算下使用475个令牌而非42,450个。Rift识别出49个混淆簇,包括在引导生成卡片中的四个高风险簇。对119个LLM生成描述进行探索性比较,消除了观察到的零距离簇,但表明混合卡片生成机制可能降低R@5。网关测量在十次试验中相对于直接stdio MCP调用增加了5毫秒平均延迟(0.8%)。Cartograph与代码执行方法互补:它控制哪些工具描述被呈现,并记录每次查询用于排名的描述来源。

英文摘要

The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog traversal to $O(k)$ progressive disclosure. Cartograph combines three mechanisms: (1) operator-attested capability cards, Ed25519-signed descriptions generated under the deploying operator's control rather than ranked publisher copy; (2) Rift, a three-layer confusable-cluster analysis comprising density clustering, query-margin analysis, and token diagnosis; and (3) two-stage retrieval, which ranks servers before tools. On a 22-server, 374-tool deployment, Cartograph exposes three proxy tools instead of 374 definitions. A 49-query author-constructed benchmark yields R@5 of 0.816, compared with 0.592 for a Jaccard keyword baseline, while a measured top-5 discovery exchange uses 475 tokens rather than 42,450 under the stated full-catalog accounting. Rift identifies 49 confusable clusters, including four HIGH-risk clusters in bootstrap-generated cards. An exploratory comparison of 119 LLM-generated descriptions removes the observed zero-distance cluster but shows that mixing card-generation regimes can reduce R@5. Gateway measurements over ten trials add 5ms mean latency (0.8%) relative to direct stdio MCP calls. Cartograph is complementary to code-execution approaches: it controls which tool descriptions are surfaced and records the provenance of the descriptions used for ranking for each query.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑