arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28033cs.AI

DataClawEval:面向真实工业场景的数据工程智能体基准测试

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness

Debin Meng, Jiaming Yang, Zefang Zong, Tengyue Xu, Haining Xie, Yang Li, Peng Chen

首次发表
浏览论文内容

中文总结 AI 辅助

DataClawEval是首个针对真实工业数据工程场景的智能体基准,含100项跨5种引擎的任务,用确定性脚本评分,评估16个前沿智能体发现最强模型仅得74.9分,自主数据工程仍是挑战。

中文摘要 AI 辅助

大型语言模型(LLMs)及基于LLM的智能体正越来越多地被部署用于自动化复杂工作流,有望彻底变革数据管理与处理。然而,现有基准测试主要聚焦于简化的Text-to-SQL翻译或数据分析,却基本未涉及关键且复杂的端到端数据工程领域。为填补这一空白,我们推出DataClawEval,这是首个专门用于评估真实世界数据工程场景中自主智能体端到端任务完成能力的综合基准测试。该基准基于专业企业数据工程师编写的生产级代码构建,包含100项严谨的端到端任务,覆盖PySpark、MySQL、HiveSQL、PrestoSQL/Trino和FlinkSQL五种执行引擎。与非确定性的LLM作为评判者的评分方式不同,每项任务在特定的隔离沙箱中执行,并通过确定性的基于规则的脚本进行评分。对16个前沿智能体的评估揭示了关键局限性:表现最强的模型仅取得74.9的总分,且没有单一模型占据主导,因为每个模型在不同引擎上表现出色,这凸显了严格的领域专业化而非全能能力。因此,自主数据工程仍是一项艰巨且未解决的挑战。我们在该httpsURL发布了我们的数据集、容器化环境以及确定性评估脚本。

英文摘要

Large language models (LLMs) and LLM-based agents are increasingly being deployed to automate complex workflows, promising to revolutionize data management and processing. However, existing benchmarks predominantly focus on simplified Text-to-SQL translation or data analysis, leaving the critical and complex domain of end-to-end data engineering largely unexplored. To bridge this gap, we introduce DataClawEval, the first comprehensive benchmark designed specifically to evaluate the end-to-end task completion capabilities of autonomous agents in real-world data engineering scenarios. Built upon production-grade code authored by professional enterprise data engineers, it comprises 100 rigorous, end-to-end tasks spanning five execution engines: PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL. Rather than non-deterministic LLM-as-a-judge scoring, each task is executed within a case-specific, isolated sandbox and graded by deterministic, rule-based scripts. Evaluating 16 frontier agents exposes critical limitations: The strongest model attains only 74.9 overall, and no single model dominates, as each excels on a different engine, revealing strict domain specialization rather than omnipotent proficiency. Thus, autonomous data engineering remains a formidable, unresolved challenge. We release our dataset, containerized environments, and deterministic evaluation scripts at https://github.com/Dicemy/DataClawEval/tree/master

发表机构

  • Tencent(腾讯)
  • Xidian University(西安电子科技大学)
  • South China Normal University(华南师范大学)

机构由 AI 辅助整理,请以论文原文为准。

↑