AI 中文总结
本研究在BFCL v4上对比14种模型的程序化工具调用与JSON工具调用,发现程序化工具调用在多数场景下性能更优且鲁棒性更强,是可行的替代方案。
AI 中文摘要
工具使用将大语言模型(LLM)转化为能在其训练数据之外执行动作的智能体;对于具备代码能力的模型而言,程序化工具调用(PTC)进一步扩展了能力,它用能自然链式调用和并行化的脚本替代了僵化的JSON调用。然而,在真实任务条件下,针对现有基准测试集、覆盖当前及前代模型的代码类工具系统评估尚未开展。本研究在BFCL v4基准测试集上,对14种语言模型的程序化工具调用(PTC)与原生JSON工具调用进行了实证对比。在程序化工具调用范式中,工具以类型化Python存根的形式呈现,模型通过代码调用这些存根,执行过程与结果处理在单个智能体回合内完成。在BFCL v4上,14种模型中有11种的程序化工具调用性能与原生JSON工具调用相当或更优,其中GPT-5.6家族较JSON工具调用基线实现了10.6%的提升;在并行扇出场景下,14种模型中有13种的程序化工具调用性能与基线相当或更优,且在上下文衰减条件下表现稳定,而基线平均性能下降2.3%。研究结果表明,程序化工具调用是JSON工具调用可行且鲁棒的替代方案,其性能随模型发布代次的能力变化而变化。
英文摘要
Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by replacing rigid JSON calls with scripts that chain and parallelize naturally. However, a systematic evaluation of tools as code on an established benchmark across current and prior model generations under real-world task conditions has not been conducted. In this work, we empirically compare programmatic tool calling (PTC) to native JSON tool calling across 14 language models on BFCL v4. In the programmatic tool calling paradigm, tools are exposed as typed Python stubs that the model invokes through code, with execution and results handled in a single agent turn. Programmatic tool calling matches or exceeds native JSON tool calling in 11 of 14 models on BFCL v4, with the GPT-5.6 family achieving a 10.6% improvement over the JSON tool calling baseline. Further, it matches or outperforms baseline in 13 of 14 models under parallel fan-out, and holds stable under context rot conditions where baseline degrades 2.3% on average. Our results demonstrate that programmatic tool calling is a viable and robust alternative to JSON tool calling, with performance tracking model capability across release generations.