发表机构
Hunyuan Team, Tencent; Institute for AI Industry Research, Tsinghua University; School of Computer Science and Engineering, Southeast University(腾讯混元团队; 清华大学人工智能产业研究院; 东南大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型多步工具使用能力,介绍E-Bench基准测试,通过解耦环境与任务合成构建可扩展可控测试平台,对11个前沿模型测试发现多步工具使用仍具挑战,即便结合代码执行可靠性也不高。
AI 中文摘要
大语言模型(LLMs)越来越多地被用作在多个步骤中与有状态环境交互的智能体,即多步工具使用。现有基准测试虽推动了工具使用智能体评估,但存在局限。我们引入E-Bench,这是一个完全合成的基准测试,涵盖三个产品领域的323个状态改变任务。它将环境合成与任务合成解耦,通过图引导数据库填充构建环境,生成器-求解器不对称创建任务。通过数据库状态差异确定性评分,在环境层面可控,任务层面可扩展。对11个前沿LLMs测试表明多步工具使用仍具挑战,最强模型的Pass^3低于60%,即便使用E-Bench-Code扩展中的代码执行,可靠性(Pass^3)仍低于70%。
英文摘要
Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes. We refer to this capability as multi-step tool use. Existing benchmarks have advanced tool-use agent evaluation, but often focus on isolated API calls, short trajectories, or settings that are difficult to scale or control. We introduce E-Bench, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting. E-Bench decouples environment synthesis from task synthesis: graph-guided database filling builds reusable, orphan-free product environments, while generator-solver asymmetry creates tasks with both an information gap and a tool gap, requiring agents to discover hidden data and compose multiple tool calls before changing state. Outcomes are graded deterministically by database-state diffs. Since both environments and tasks are synthetic, E-Bench is controllable at the environment level and scalable at the task level. Benchmarking 11 cutting-edge LLMs shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code extension, reliability (Pass^3) remains below 70%.
Comments29 pages, 14 figures, 6 tables