arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04611cs.AI

$τ^τ$-Bench:用于端到端、真实智能体构建的环境

$τ^τ$-Bench: An Environment for End-To-End, Realistic Agent Construction

Quan Shi, Keshav Dhandhania, Karthik Narasimhan, Victor Barres

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出$τ^τ$-bench基准测试,将智能体构建作为任务,测试编码智能体在真实场景下交付客户服务智能体的能力,发现当前最强配置通过率仅23.9%,该基准可衡量协作式智能体构建效果。

中文摘要 AI 辅助

大语言模型(LLM)智能体正迅速成为生产级软件,被部署用于处理客户服务、裁决纠纷和运营内部系统。值得注意的是,构建这类智能体的工作正越来越多地交由编码智能体完成,但现有基准几乎未涉及AI系统能否在真实客户交互场景下交付可用智能体的问题。我们推出$τ^τ$-bench(发音为hyper-tau-bench),这是一项将智能体构建作为核心任务的基准测试。开发智能体将获得企业实际保留的记录、提出需求的客户、操作必须调用的生产API、可继承的代码库,以及服务成本和模型的限制——这些均与真实交互场景的起始条件一致。开发智能体需据此交付完整的客户服务智能体,评估方式为将该智能体部署到预留的模拟用户中进行测试。在涵盖四个领域的53项任务中,最强配置Claude Opus 5在Claude Code下的评估模拟通过率仅为23.9%,而专家编写的参考基准得分为82.2%。失败情况与人类智能体开发者遇到的问题类似:模型仅发出浅层查询,而非对记录进行深度理解;几乎不与客户沟通;在智能体架构和服务成本上尝试不足,直接交付首个可运行的设计。我们希望$τ^τ$-bench能将协作式智能体构建工作转化为编码智能体可衡量的目标。

英文摘要

LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce $τ^τ$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for $τ^τ$-bench to turn the work of cooperative agent building into a measurable target for coding agents.

发表机构

  • Princeton University(普林斯顿大学)
  • Sierra(塞拉公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑