Toollery:将LLM智能体扩展到数千技能与工具
Toollery: Scaling LLM Agents to Thousands of Skills and Tools
浏览论文内容
中文总结 AI 辅助
针对LLM智能体面临大量技能工具时全库提示成本高、选择可靠性低的问题,提出无需训练的候选压缩框架Toollery,通过生成意图查询构建检索索引,将候选集压缩至top-k,在多个基准上提升召回率并保持准确率。
中文摘要 AI 辅助
随着LLM智能体暴露于数百至数万种技能、工具和API函数,全库提示变得昂贵、缓慢且可靠性降低:每个新增候选都会增加提示词token和延迟,而更长的候选列表会为LLM选择引入更多干扰项。我们提出\textbf{Toollery},一个无需训练的候选压缩框架,用于可扩展的LLM技能/工具选择。遵循既有的文档侧查询扩展方法,Toollery从每个技能/工具规范中生成用户意图查询,并构建一个检索索引,在最终LLM决策之前将真实用户请求映射到紧凑的候选集。通过将高层技能和原子工具视为可选择的capabilities,Toollery可应用于技能库和工具注册表。我们在约79K能力的SkillRouter基准、包含超过440个原子工具的BFCL-V4以及3,396个专有智能座舱请求(覆盖220个工具)上评估Toollery。在这些设置中,Toollery将在线选择限制在紧凑的top-$k$候选集内,并提高了相对于普通规范检索的召回率。在固定的top-10预算下,Toollery在座舱数据集上提升了端到端选择性能,并在BFCL-V4上保持了相当的AST准确率。这些结果支持Toollery作为大型且不断演进的智能体能力库的实用候选压缩框架,同时表明质量和成本收益取决于工作负载覆盖率和提供方缓存。
英文摘要
As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for LLM selection. We present \textbf{Toollery}, a training-free candidate-compression framework for scalable LLM skill/tool selection. Following established document-side query expansion, Toollery generates user-intent queries from each skill/tool specification and builds a retrieval index that maps real user requests to compact candidate sets before final LLM decision-making. By treating high-level skills and atomic tools as selectable capabilities, Toollery can be applied to both skill libraries and tool registries. We evaluate Toollery on the roughly 79K-capability SkillRouter benchmark, BFCL-V4 with over 440 atomic tools, and 3,396 proprietary smart-cockpit requests over 220 tools. Across these settings, Toollery keeps online selection bounded to a compact top-$k$ candidate set and improves recall over ordinary specification retrieval. At a fixed top-10 budget, Toollery improves end-to-end selection on the cockpit dataset, and maintains comparable AST Accuracy on BFCL-V4. These results support Toollery as a practical candidate-compression framework for large and evolving agent capability libraries, while showing that quality and cost gains depend on workload coverage and provider caching.
发表机构
- Laboratories, Huawei(华为2012实验室)
机构由 AI 辅助整理,请以论文原文为准。