从SQL生成到工具选择:面向MCP服务器的领域导向模式
From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers
浏览论文内容
中文总结 AI 辅助
该研究提出面向MCP服务器的领域导向工具模式,通过用意图分类替代SQL合成降低模型需求,在Sakila数据库基准测试中,垂直领域工具包表现最优,最小模型成本大幅降低且性能达标。
中文摘要 AI 辅助
基于大语言模型(LLM)构建的智能体正通过模型上下文协议(MCP)越来越多地访问企业数据,许多MCP数据库服务器通过暴露单一通用SQL执行工具来最大化灵活性。本文提出面向领域的工具模式:模型不在查询时生成SQL,而是从少量与领域对齐的工具中进行选择,这些工具的参数化查询在服务器端封装了架构导航、连接和业务规则。我们围绕三个架构不变量将该模式形式化,并引入模型降级概念,即观察到用意图分类取代SQL合成可降低处理常规请求所需的模型层级。作为参考实现,我们推出开源框架MCP Blueprint,其中领域工具以声明方式定义为YAML元数据加外部参数化SQL文件。我们通过公开可复现基准评估该模式,在Sakila数据库上针对17项面向客户的任务,比较三种MCP服务器设计——原始SQL执行、精简通用工具包和垂直领域工具包,涉及四个本地模型(3B-8B),共609个完成单元,温度为0,每个单元重复三次。垂直领域工具包的合并平均得分为0.939,而原始SQL为0.666,通用工具包为0.605;最小模型的得分从0.583提升至0.929,达到或超过所有更大规模配置,同时将每个正确答案的成本降低一个数量级。所有测试代码、提示、标准答案、冻结工具包及每个单元结果均公开可用。
英文摘要
Agents built on Large Language Models (LLMs) increasingly reach enterprise data through the Model Context Protocol (MCP), and many MCP database servers maximize flexibility by exposing a single generic SQL execution tool. This paper proposes the Domain-Oriented Tooling Pattern: instead of generating SQL at query time, the model selects from a small set of domain-aligned tools whose parameterized queries encapsulate schema navigation, joins and business rules on the server side. We formalize the pattern around three architectural invariants and introduce Model Demotion, the observation that replacing SQL synthesis with intent classification lowers the model tier required to serve routine requests. As a reference implementation we present MCP Blueprint, an open-source framework in which domain tools are defined declaratively as YAML metadata plus external parameterized SQL files. We evaluate the pattern with a public reproducibility benchmark comparing three MCP server designs - raw SQL execution, a thin generic tool pack, and a verticalized domain pack - on four local models (3B-8B) across seventeen customer-facing tasks over the Sakila database (609 completed cells; temperature 0; three repetitions per cell). The verticalized pack reaches a pooled mean score of 0.939 versus 0.666 for raw SQL and 0.605 for the generic pack; the smallest model improves from 0.583 to 0.929, matching or exceeding every larger configuration while cutting cost per correct answer by an order of magnitude. All harness code, prompts, gold answers, frozen packs and per-cell results are publicly available.