StartupBench:基于市场验证的端到端工作流程的通用智能体基准测试
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
浏览论文内容
中文总结 AI 辅助
本文提出StartupBench基准测试,基于市场验证的AI初创产品工作流程评估通用智能体,发现当前最强模型仅完成约30%任务,该基准可衡量智能体完成现实用户任务的进展。
中文摘要 AI 辅助
大型语言模型(LLMs)和智能体的最新进展大幅提升了AI系统执行复杂任务的能力。然而,现有基准测试大多依赖研究人员选定的任务,无法确定这类进展是否延伸到现实世界用户对AI的实际需求工作中。我们推出StartupBench,这是一个基于市场验证的AI初创产品的端到端智能体基准测试。我们没有根据对智能体有用能力的预先定义假设来定义任务,而是系统研究已被采用的AI产品及其产品工作流程和用户,以识别AI在不同专业领域已确立实际需求的现实世界任务。我们将这些工作流程转化为完整的、面向可交付成果的任务,并通过捕捉其复杂要求的细粒度评分规则对其进行评估。在统一智能体框架下评估的代表性模型中,即使是最强的模型也仅成功完成了StartupBench的约30%,尽管在许多任务上取得了实质性的部分进展。进一步分析表明,复杂指令遵循和领域特定专业知识等方面是主要失败来源。我们的结果表明,许多市场验证的工作流程仍然超出了当前通用智能体的可靠能力范围,确立了StartupBench作为衡量现实世界用户任务端到端完成进展的经验指标。
英文摘要
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
发表机构
- ByteDance Seed(字节跳动 Seed)
- Nanjing University(南京大学)
- M-A-P
机构由 AI 辅助整理,请以论文原文为准。