发表机构
School of Mathematical Sciences, Beijing University of Posts and Telecommunications; School of Mathematics and Statistics, Chongqing University; School of Science, Beijing Forestry University(北京邮电大学数学科学学院; 重庆大学数学与统计学院; 北京林业大学理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM工具使用中渐进式披露与提示词缓存的权衡,提出CacheRouter双路径路由架构,通过主模型隔离与工具路由通道设计提升缓存命中率,降低输入成本。
AI 中文摘要
大语言模型(LLM)系统中的工具使用面临结构性权衡:渐进式披露仅展示当前任务相关工具以保持提示词短小,而提示词缓存则要求请求前缀在多次调用间保持固定,对可见工具列表的任何更改都会使缓存前缀失效。本文将该权衡视为请求架构问题,提出一种双路径路由设计,将工具选择与工具交付分配至独立通道:主模型始终可见一组小型固定核心工具,因此其请求头部在多次调用间保持不变;所有其他工具通过独立路由通道访问,其中路由器子模型会搜索完整工具列表、选择一个工具、执行并返回结果。工具注册可从源代码自动完成并支持运行时更新,因此工具集可在不修改主模型请求前缀的情况下扩展。该设计推广了渐进式披露:能力通过路由通道披露,主模型前缀保持稳定。原型实现在55个功能查询和30轮对话中进行测试,令牌级缓存命中率分别达到90.99%和95.2%,基于DeepSeek的定价,输入成本降至无缓存基准的约12.0%和8.0%,其中缓存命中的输入令牌成本约为缓存未命中令牌的1/30。
英文摘要
Tool use in LLM systems faces a structural trade-off. Progressive disclosure keeps the prompt small by showing only the tools relevant to the current task, while prompt caching rewards a request prefix that stays fixed across calls; every change to the visible tool list invalidates the cached prefix. This paper treats the trade-off as a problem of request architecture and proposes a dual-path routing design that assigns tool selection and tool delivery to separate channels. The main model always sees a small, fixed set of core tools, so the head of its request is unchanged across calls; all other tools are reached through an independent routing channel, in which a router sub-model searches the full tool list, selects one tool, executes it, and returns the result. Tool registration is automated from source code and supports runtime updates, so the tool set can grow without modifying the main model's request prefix. The design generalizes progressive disclosure: capabilities are disclosed through the routing channel, and the main model's prefix stays stable. A prototype implementation was exercised on 55 functional queries and a 30-turn dialogue; token-level cache hit rates reached 90.99% and 95.2%, cutting input cost to about 12.0% and 8.0% of a no-cache baseline under DeepSeek's pricing, where cache-hit input tokens cost roughly 1/30 of cache-miss tokens.