arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DataSpace:针对异构工作空间可验证分析的数据智能体基准测试

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo

arXiv 2608.03451首次发表:更新:

AI 中文总结

该研究推出DataSpace基准,用于评估异构工作空间下数据智能体的可验证分析能力,包含多类型任务与数据,实验揭示了智能体框架、多模态集成等对准确率的影响,为提升数据智能体可靠性指明了方向。

AI 中文摘要

数据智能体支持对组织工作空间进行自然语言分析,相关证据可能分散在数据库、结构化文件、长文档和多媒体中。现有基准大多将结构化查询、检索或开放式分析隔离开来,未能统一处理异构证据发现、完整表格输出和确定性评估。我们推出DataSpace,这是一个数据智能体从任务本地异构工作空间生成可验证表格结果的基准。它包含410项跨语言任务和7439个工件,总大小15.01 GB,涵盖CSV、JSON、SQLite、Markdown、PDF和视频格式。DataSpace同时是2026年KDD Cup“复杂数据分析数据智能体”竞赛的官方评估基准。每个智能体仅接收问题和工作空间,返回完整的请求表格结果。我们使用DataSpace-Builder构建DataSpace,这是一个基于执行的框架,包含跨语言转换、约束感知关系采样、模态路由与工件渲染,以及11位领域专家的人工审核和任务修复。确定性评估器执行与表头无关的列对齐、类型和精度感知的归一化,以及与顺序相关的行比较。在6个最新发布的前沿多模态模型和5个广泛使用的智能体框架上,最佳准确率达到66.34%,而在主干模型固定的情况下,框架选择造成了15.36个百分点的差异。多模态证据集成和连接操作在所有6个主干模型上均持续降低准确率。这些结果表明DataSpace尚未达到饱和,并确定了提高数据智能体可靠性的关键挑战。

英文摘要

Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic evaluation insufficiently unified. We introduce DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces. It contains 410 cross-language tasks and 7,439 artifacts totaling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. DataSpace also served as the official evaluation benchmark for the KDD Cup 2026 Data Agents for Complex Data Analysis competition. Each agent receives only a question and workspace and returns the complete requested tabular result. We construct DataSpace with DataSpace-Builder, an execution-grounded framework comprising cross-language transformation, constraint-aware relational sampling, modality routing and artifact rendering, and human review and task repair by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison. Across six recently released frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%, while harness choice creates a 15.36-point spread with the backbone fixed. Multimodal evidence integration and joins consistently reduce accuracy across all six backbones. These results show that DataSpace remains unsaturated and identify key challenges for improving data-agent reliability.

Comments8 pages of main text, 7 figures, with a supplementary appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑