arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AlloBench:测量大语言模型智能体中的在线工具分配能力

AllocBench: Measuring Online Tool Allocation Capability in LLM Agents

Daniel Wang, Andrew Xu

arXiv 2607.23332首次发表:更新:

发表机构

Sea12 Technologies; Yale University(Sea12科技公司; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究测试大语言模型智能体在固定预算下的在线工具分配能力,通过配对基准测试发现前沿模型在抽象框架中表现近乎最优,但在脚本编写中失败,确定了各模型失败模式,还表明开源Qwen模型在抽象分配上有推广,在线工具分配是重要能力边界。

AI 中文摘要

创建可复用工具是一种投资:智能体现在支付固定成本以换取未来复用的潜力。因此,用户应更喜欢创建少量高度可复用工具的智能体,而非许多一次性工具。我们引入了一个配对基准测试,在两种情境下(基于抽象文本的表述和代码构建任务)测试大语言模型智能体在固定预算下是否表现出有意识的分配行为。我们发现,测试的每个前沿模型(Claude Haiku、Claude Opus、GPT - 5.4 - mini和GPT - 5.6 Sol)在抽象框架中表现近乎最优,但无法将此能力转移到脚本编写中。通过进一步实验,我们确定了每个模型的特定失败模式。值得注意的是,前三个模型即使在不评估脚本时也失败,而GPT - 5.6 Sol在较弱操作下保持选择性,仅在完整构建时崩溃。此外,为抽象分配进行策略训练的开源Qwen模型在保留的词汇变体中推广了此能力,但在脚本分配上没有改进。这些结果表明,即使对于现代前沿模型,在线工具分配也是一个重要的能力边界。

英文摘要

Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer an agent that creates a small number of highly reusable tools, rather than many one-offs. We introduce a paired benchmark that tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in two contexts: an abstract text-based formulation and a code-construction task. We find that every frontier model we test---Claude Haiku, Claude Opus, GPT-5.4-mini, and GPT-5.6 Sol---acts near-optimally in the abstract framing but fails to transfer this ability to script-writing. Through further experiments, we identify the particular failure modes for each model. Notably, the first three models fail even when the scripts are not evaluated, while GPT-5.6 Sol stays selective under that weaker manipulation and collapses only at full construction. Furthermore, an open-source Qwen model policy-trained for abstract allocation generalizes this ability across held-out lexical variations, but sees no improvement at script allocation. Together, these results establish online tool allocation as a significant capability boundary, even for modern frontier models.

Comments24 pages, 6 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑