arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17800cs.AI

StartupBench:基于市场验证的端到端工作流程的通用智能体基准测试

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi L… 展开作者

Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出StartupBench基准测试,基于市场验证的AI初创产品工作流程评估通用智能体,发现当前最强模型仅完成约30%任务,该基准可衡量智能体完成现实用户任务的进展。

中文摘要 AI 辅助

大型语言模型(LLMs)和智能体的最新进展大幅提升了AI系统执行复杂任务的能力。然而,现有基准测试大多依赖研究人员选定的任务,无法确定这类进展是否延伸到现实世界用户对AI的实际需求工作中。我们推出StartupBench,这是一个基于市场验证的AI初创产品的端到端智能体基准测试。我们没有根据对智能体有用能力的预先定义假设来定义任务,而是系统研究已被采用的AI产品及其产品工作流程和用户,以识别AI在不同专业领域已确立实际需求的现实世界任务。我们将这些工作流程转化为完整的、面向可交付成果的任务,并通过捕捉其复杂要求的细粒度评分规则对其进行评估。在统一智能体框架下评估的代表性模型中,即使是最强的模型也仅成功完成了StartupBench的约30%,尽管在许多任务上取得了实质性的部分进展。进一步分析表明,复杂指令遵循和领域特定专业知识等方面是主要失败来源。我们的结果表明,许多市场验证的工作流程仍然超出了当前通用智能体的可靠能力范围,确立了StartupBench作为衡量现实世界用户任务端到端完成进展的经验指标。

英文摘要

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

发表机构

  • ByteDance Seed(字节跳动 Seed)
  • Nanjing University(南京大学)
  • M-A-P

机构由 AI 辅助整理,请以论文原文为准。

↑