arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18171cs.AIcs.LG

循环语言模型可提升组合式工具调用能力

Looped Language Models Improve Compositional Tool Calling

Andrei Cristian Popescu, Haitz Sáez de Ocáriz Borde, Pietro Liò

首次发表
浏览论文内容

中文总结 AI 辅助

该研究在组合式工具调用场景下,通过在多个数据集上对比循环与非循环语言模型,发现循环计算可提升组合式工具使用性能,自适应推理能实现更优计算-性能权衡,为智能体系统提供了有前景的架构。

中文摘要 AI 辅助

循环语言模型在推理基准测试中已展现出良好效果,但其在智能体工具使用方面的潜力仍未得到充分探索。我们在组合式工具调用场景中研究该问题,该场景要求模型协调多个API调用、维护中间状态并保留工具交互间的依赖关系。我们在API-Bank、BFCL和NESTful数据集上评估原生及改造后的循环语言模型,比较在匹配的监督微调方案下训练的循环与非循环模型,并在推理时改变循环深度。受控实验表明,循环计算通常有益于组合式及依赖感知的工具使用,而在独立API调用上的增益较小且依赖于模型;多步骤工具使用的准确率通常随循环深度增加而提升,然而自适应推理通过仅在需要时分配额外计算,实现了更优的计算-性能权衡。我们的结果表明,循环语言模型是需要可靠规划、协调及执行组合式工具使用工作流的智能体系统的有前景架构。

英文摘要

Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use remains largely unexplored. We study this question in compositional tool-calling settings, where models must coordinate multiple API calls, maintain intermediate state, and preserve dependencies across tool interactions. We evaluate native and retrofitted looped language models on API-Bank, BFCL, and NESTful, comparing looped and non-looped models trained under matched supervised fine-tuning recipes and varying recurrent depth at inference time. In controlled experiments, recurrent computation generally benefits compositional and dependency-aware tool use, while providing smaller and more model-dependent gains on isolated API invocation. Accuracy on multi-step tool use generally increases with recurrent depth; adaptive inference, however, achieves a more favorable compute-performance trade-off by allocating additional computation only when needed. Our results suggest that looped language models are a promising architecture for agentic systems that require reliable planning, coordination, and execution of compositional tool use workflows.

发表机构

  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑