arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02122cs.CLcs.AIcs.DB

Argo-Bench:评估企业级工作流中的数据智能体

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

首次发表
浏览论文内容

中文总结 AI 辅助

Argo-Bench是一个评估框架,通过模拟纽约市外卖平台并构建企业级数据仓库,测试数据智能体在复杂工作流中的推理与行动能力,超越传统文本到SQL基准。

中文摘要 AI 辅助

真实世界中的企业数据科学与分析工作流需要对数十张表进行推理、执行统计分析并基于结果采取行动。已有的文本到SQL基准仅评估查询生成,且审计发现其答案键经常出错。由于真实的企业数据仓库过于敏感而无法发布,这些基准基于公共数据集构建,其中每个业务事件都适合放在单张表中。我们引入了Argo-Bench,一个包含210个数据科学与分析任务的评估框架。利用公共数据、同行评审的行业文献和监管文件,我们在纽约市模拟了一个真实规模的外卖平台,2024年有8100万订单,具有基于现实的经济学、欺诈模式和市场竞争激励。我们将这个模拟世界导出到一个包含235张表和75亿行的ERP数据仓库,其模式仿照Oracle E-Business Suite。模拟器的真实状态对智能体所见的仓库是隐藏的,因此任务要求智能体在采取行动之前通过导航仓库来重建事实。Argo-Bench超越了文本到SQL:智能体执行诸如封禁欺诈账户、分配骑手激励预算或发放补发工资等操作,评分器根据这些操作在模拟器中的后果进行评分。每个任务都有一个可执行的参考解决方案,证明仅使用仓库即可解决。在14个前沿和开放权重模型中,最强的模型仅在34.8%的任务上得分达到95或更高,平均得分为59.5分。我们希望Argo-Bench能推动智能体在真实数据环境中理解、导航和行动方面的进步。

英文摘要

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.

发表机构

  • TextQL

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑