循环语言模型可提升组合式工具调用能力
Looped Language Models Improve Compositional Tool Calling
浏览论文内容
中文总结 AI 辅助
该研究在组合式工具调用场景下,通过在多个数据集上对比循环与非循环语言模型,发现循环计算可提升组合式工具使用性能,自适应推理能实现更优计算-性能权衡,为智能体系统提供了有前景的架构。
中文摘要 AI 辅助
循环语言模型在推理基准测试中已展现出良好效果,但其在智能体工具使用方面的潜力仍未得到充分探索。我们在组合式工具调用场景中研究该问题,该场景要求模型协调多个API调用、维护中间状态并保留工具交互间的依赖关系。我们在API-Bank、BFCL和NESTful数据集上评估原生及改造后的循环语言模型,比较在匹配的监督微调方案下训练的循环与非循环模型,并在推理时改变循环深度。受控实验表明,循环计算通常有益于组合式及依赖感知的工具使用,而在独立API调用上的增益较小且依赖于模型;多步骤工具使用的准确率通常随循环深度增加而提升,然而自适应推理通过仅在需要时分配额外计算,实现了更优的计算-性能权衡。我们的结果表明,循环语言模型是需要可靠规划、协调及执行组合式工具使用工作流的智能体系统的有前景架构。
英文摘要
Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use remains largely unexplored. We study this question in compositional tool-calling settings, where models must coordinate multiple API calls, maintain intermediate state, and preserve dependencies across tool interactions. We evaluate native and retrofitted looped language models on API-Bank, BFCL, and NESTful, comparing looped and non-looped models trained under matched supervised fine-tuning recipes and varying recurrent depth at inference time. In controlled experiments, recurrent computation generally benefits compositional and dependency-aware tool use, while providing smaller and more model-dependent gains on isolated API invocation. Accuracy on multi-step tool use generally increases with recurrent depth; adaptive inference, however, achieves a more favorable compute-performance trade-off by allocating additional computation only when needed. Our results suggest that looped language models are a promising architecture for agentic systems that require reliable planning, coordination, and execution of compositional tool use workflows.
发表机构
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。