arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19741cs.CLcs.DB

单次成功并非可靠:Thinkingbox——面向有状态业务工作流智能体的沙箱与基准测试集

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

Zhuochun Li, Youngmin Ko, Ali Keramati, Tuhin Kundu, Liang-Chun Tsai, Nicola Ferri, Mirco Milletari, Jiaxiang Liu, Susana Palmaz Lopez Pelaez, Yuepeng Wang, Vad… 展开作者

Zhuochun Li, Youngmin Ko, Ali Keramati, Tuhin Kundu, Liang-Chun Tsai, Nicola Ferri, Mirco Milletari, Jiaxiang Liu, Susana Palmaz Lopez Pelaez, Yuepeng Wang, Vadim Smolyakov, Xiang Jiang, Kjartan Olafsson, Tommy Guy

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员推出面向有状态业务工作流智能体的沙箱Thinkingbox及基准测试集Thinkingbox-bench,发现现有智能体在该基准上可靠完成任务的表现远逊于单次成功表现,凸显了端到端任务可靠评估的必要性。

中文摘要 AI 辅助

近期的智能体基准测试日益将评估建立在可执行环境中,涵盖代码修复、网页导航、应用程序API及函数调用等场景。然而,完成代码之外的重要工作,仅生成合理响应或有效工具调用是不够的:智能体必须多轮收集缺失信息、遵循领域策略、协调依赖工具,并实现正确的持久状态转换且无附带影响。本文提出Thinkingbox,这是一个用于工具-智能体-用户交互的沙箱,提供与MCP兼容的隔离工具会话、完整执行轨迹以及针对终端后端状态的结果评估。基于该沙箱构建的Thinkingbox-bench包含507个受策略约束的工作流,覆盖零售、酒店、汽车保险、新银行内部IT及咨询IT/HR支持等众多场景。每次尝试均通过特定任务的可执行检查进行评估,该检查接受有效轨迹,拒绝错误、缺失或多余的影响;指定任务还会检查最终响应的所需属性。在专有模型和开放权重模型中,最强模型达到65.36%的pass@1,但仅为25.25%的pass^20。此外,许多失败的试验显示终止干净且状态变更操作有效,这表明响应或工具调用层面的信号并非端到端任务完成的清晰代理。Thinkingbox-bench揭示了偶尔找到成功轨迹与可靠完成有状态业务任务之间存在巨大差距。我们发布Thinkingbox和Thinkingbox-Bench:此httpsURL

英文摘要

Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across business scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Our experiments reveal that even the strongest proprietary and open-weight models show steep reliability drops: Claude Opus 5 falls from 66.50% pass@1 to 47.53% pass^20, and Kimi-K3 from 57.37% pass@1 to 17.60% pass^20. Moreover, many failed trials terminate cleanly after valid state-changing actions, so response- or tool-call-level signals poorly proxy end-to-end completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox (https://github.com/microsoft/thinkingbox) and Thinkingbox-bench (https://github.com/microsoft/thinkingbox-data).

发表机构

  • University of Pittsburgh(匹兹堡大学)
  • Northwestern University(西北大学)
  • University of California Irvine(加州大学欧文分校)
  • Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

↑