arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MCP中的大语言模型很重要:测量由大语言模型驱动的低效资源利用

LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

Minhan Cho, Soyoung Park, Kihyeon Jeong, Byeongkyu Jeon, Daejin Choi, Jinyoung Han

arXiv 2608.08467首次发表:更新:

发表机构

Sungkyunkwan University; National Assembly Research Service; AlphaBridge; Ewha Womans University(成均馆大学; 国会研究服务处; 阿尔法桥公司; 梨花女子大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对24个LLM开展54000次试验,发现MCP中服务器嵌入的参考数据会因LLM偏好导致低效资源利用,提出需将服务器指令置于客户端LLM工具选择之前的改进方向。

AI 中文摘要

模型上下文协议(MCP)标准化了服务器向大语言模型(LLM)暴露数据和工具的方式。一种常见的服务器设计会将常用参考数据(如标识符查找表)直接嵌入服务器指令中,即服务器传递给主机应用程序的系统提示文本。当查询涉及嵌入表中的条目时,模型可直接对其进行操作,而非通过搜索工具重新发现相同信息。我们测试客户端LLM是否实际使用此类嵌入指令的数据,报告了一项在生产级法律信息MCP服务器上对24个LLM(9个Claude、6个Gemini、9个GPT)开展的54000次试验研究。移除竞争搜索工具的诊断条件显示,失败主要源于行为偏好而非能力缺失:搜索不可用时,24个模型中有23个能可靠读取嵌入数据(命中率至少98%);仅存在搜索工具时,9个模型的命中率降至15%以下。对三种指令级干预措施的2^3析因分析显示存在强交互效应:结合全部三种干预措施可使24个模型中的20个命中率恢复至至少86%,但单独干预可能对特定模型家族产生反效果。因此,每服务器的提示工程只是权宜之计而非解决方案;我们认为MCP主机应用程序应提供明确机制,将服务器指令置于客户端LLM决策的工具选择之前。

英文摘要

The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequently used reference data, such as identifier lookup tables, directly in the server instructions: the system-prompt text a server hands to the host application. When a query concerns an entry of the embedded table, the model can act on it immediately instead of re-discovering the same information through a search tool. We test whether client LLMs actually consume such instruction-embedded data, reporting a 54,000-trial study across 24 LLMs (9 Claude, 6 Gemini, 9 GPT) on a production legal-information MCP server. A diagnostic condition that removes the competing search tool shows that failures are dominated by behavioral preference rather than missing capability. With search unavailable, 23 of 24 models read the embedded data reliably (hit ratio at least 98%); with a search tool merely present, 9 models drop below 15%. A 2^3 factorial analysis of three instruction-level interventions reveals strong interaction effects: combining all three restores at least 86% for 20 of 24 models, but individual interventions can backfire for specific model families. Per-server prompt engineering is therefore a workaround rather than a fix; we argue that MCP host applications should provide an explicit mechanism that places server instructions ahead of tool selection in the client LLM's deliberation.

Comments4 pages, 1 table. Accepted at the AgentSearch Workshop at SIGIR 2026, Melbourne, Australia (non-archival). Code and data: https://github.com/rabqatab/llm-in-mcp-matters

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑