AI 中文总结
本研究提出AgentWeave,一种推理前确定性路由层,可缩小工具丰富语言模型的函数调用候选集,在BFCL V4多函数任务上实现12.5%的原生BFCL成功,同时减少工具数量、输入token和延迟。
AI 中文摘要
大型语言模型越来越多地在大量工具、函数、API和专用智能体的集合上运行。随着候选动作空间扩大,函数调用模型必须处理更多模式、消耗更多提示词token,并在日益相似或不相关的选项间做出区分。我们研究一种互补的系统策略:在语言模型推理前缩小候选集,同时保持下游模型不变。我们提出AgentWeave,一种确定性推理前路由层,它利用资格、需求、能力和路由信号构建有界的模型可见动作空间。我们使用基于BFCL的路由压力协议,结合公开的MadeAgents/Hammer2.1-1.5b模型评估AgentWeave。在48个全新的BFCL V4多函数任务上,AgentWeave实现6/48(12.5%)的原生BFCL成功,而全工具、确定性随机前8和语义前8基线均实现0/48。配对成功差异为+12.5个百分点,10000次重采样配对自助法95%置信区间为+4.17至+22.92个百分点,精确McNemar检验p值为0.03125。相对于全工具暴露,AgentWeave呈现的工具减少70.18%,输入token使用量减少61.70%,平均局部模型延迟降低50.95%。该结果是刻意限定的:这是一项基于BFCL的路由压力研究,而非官方完整BFCL排行榜分数,绝对任务成功率仍然较低。不过,证据表明候选空间构建可显著影响固定模型的函数调用行为,并推动将路由作为模型推理前的独立阶段进行评估。
英文摘要
Large language models increasingly operate over large collections of tools, functions, APIs, and specialized agents. As the candidate action space grows, a function-calling model must process more schemas, consume more prompt tokens, and distinguish among increasingly similar or irrelevant alternatives. We study a complementary systems strategy: reduce the candidate set before language-model inference while leaving the downstream model unchanged. We introduce AgentWeave, a deterministic pre-inference routing layer that constructs a bounded model-visible action space using eligibility, requirement, capability, and routing signals. We evaluate AgentWeave with a frozen BFCL-derived routing-pressure protocol using the public MadeAgents/Hammer2.1-1.5b model. On 48 fresh BFCL V4 multiple-function tasks, AgentWeave achieves 6/48 (12.5%) native BFCL successes, whereas all-tools, deterministic random top-8, and semantic top-8 baselines each achieve 0/48. The paired success difference is +12.5 percentage points with a 10,000-resample paired bootstrap 95% confidence interval of +4.17 to +22.92 points and exact McNemar p=0.03125. Relative to all-tools exposure, AgentWeave presents 70.18% fewer tools, uses 61.70% fewer input tokens, and exhibits 50.95% lower mean local-model latency. The result is deliberately narrow: this is a BFCL-derived routing-pressure study rather than an official full BFCL leaderboard score, and absolute task success remains low. The evidence nevertheless shows that candidate-space construction can materially affect a fixed model's function-calling behavior and motivates evaluating routing as a distinct stage before model reasoning.
Comments12 pages, 2 figures, 6 tables. Open-source implementation and reproducibility artifacts available in the AgentWeave repository