arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

云中可扩展的大语言模型智能体工具访问

Scalable LLM Agent Tool Access in the Cloud

Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, Jing Tie, Zhou Shao, Xiaobo Xue, Xiong Xiao, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Zihao Fan, Haonan Li, Tian Pan, Xiaomin Wu, Yang Song, Xing Li, Biao Lyu, Meng Li, Haipeng Dai, Guihai Chen, Shunmin Zhu

arXiv 2607.15593首次发表:更新:

发表机构

Nanjing University; Alibaba Cloud; Fudan University; RMIT University; Politecnico di Milano; Zhejiang University(南京大学; 阿里云; 复旦大学; 皇家墨尔本理工大学; 米兰理工学院; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型智能体在云环境下工具访问问题,提出云规模网关系统,通过打破直接连接模型等整合多种功能,实现高召回率、扩展工具访问数量、提高准确性并减少时间和令牌使用量,还分享了部署经验。

AI 中文摘要

大语言模型智能体越来越依赖工具调用与外部系统交互,模型上下文协议(MCP)成为事实上的接口。但在云规模下运行MCP困难重重。工具提供方存在遗留服务难通过MCP直接调用及协议开发带来兼容性成本问题。智能体方面,可访问工具数量受限于大语言模型上下文窗口和推理开销。本文提出云规模的网关系统,打破数据平面直接连接模型,卸载遗留服务集成,整合多种功能。混合检索召回率达98%,将智能体工具访问扩展到3000多个,提高工具选择准确性,减少工具选择时间和令牌使用量,每调用开销低,扩展时稳定。最后分享了生产中部署网关系统的经验教训。

英文摘要

LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy services are not directly callable through MCP; the rapid protocol development also creates ongoing compatibility cost. On the agent side, the number of accessible tool is limited by the LLM context window and inference overhead; mounting a large tool set increases token usage and inference latency and can reduce task success rate. Moreover, for stateful MCP backends with multiple replicas, preserving session affinity increases client-side complexity. We present a cloud-scale gateway system for MCP service. It breaks the direct-connect model on the data plane and offloads legacy service integration, consolidating incompatible MCP variants, access control, tool recommendation, and session-aware routing to the gateway. Hybrid retrieval sustains 98% Top-15 recall; it scales agent tool access to 3,000+ with high tool selection accuracy, and reduces tool selection time by $8.9\times$ and token usage by $23.8\times$, with low per-call overhead, stable under scale-out. Finally, we share the lessons learned from deploying the gateway system in production.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑