arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DSAgentBench:智能体能否在真实计算机环境中自动化端到端数据科学工作流?

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince

arXiv 2608.10366首次发表:更新:

发表机构

York University; Nanyang Technological University; Salesforce AI Research(约克大学; 南洋理工大学; Salesforce AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DSAgentBench是首个评估智能体在真实计算机环境中自动化完整数据科学工作流的基准,含275项任务,实验显示现有智能体与实际需求存在显著能力差距。

AI 中文摘要

现实世界的数据科学涉及跨度较长的工作流,涵盖数据整理、探索、建模、可视化和验证,需要在真实操作系统环境中协同使用笔记本、IDE、终端、浏览器和数据库等工具。然而现有的基准测试缺乏真实计算机交互,且未评估智能体能否在现实计算环境中执行完整的端到端数据科学工作流,无法体现数据科学实践的多阶段、多工具特性。我们推出DSAgentBench,这是首个用于评估智能体能否在真实计算机环境中自动化完整数据科学工作流的基准测试。DSAgentBench包含275项涵盖整个数据科学生命周期的多样化任务,反映了实践中所需的复杂性和工具协同性。每项任务要求基于中间输出进行决策并协同使用工具,且包含确定性评估器,用于验证分析正确性、可视化输出和模型性能,而非仅验证代码执行。我们对15种闭源和开源模型开展的大量实验表明,即使是最强的智能体Claude-4.6-Sonnet也仅达到56.70%的任务成功率,而所有开源智能体的成功率均低于1%,且常在工具编排、操作系统环境适配和多步推理环节失败。这些结果揭示了当前智能体系统与真实数据科学工作流之间存在巨大的能力差距,使DSAgentBench成为开发具备环境适配性、可验证的自主数据科学智能体的基础。我们在该httpsURL发布DSAgentBench。

英文摘要

Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments. Yet existing benchmarks lack real-computer interaction and do not evaluate whether agents can execute complete end-to-end data-science workflows in realistic computing environments, failing to capture the multi-stage, multi-tool nature of data-science practice. We introduce DSAgentBench, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments. DSAgentBench contains 275 diverse tasks covering the entire data-science life-cycle, reflecting the complexity and tool coordination required in practice. Each task requires grounding decisions in intermediate outputs and coordinated tool use, and includes a deterministic evaluator that verifies analytical correctness, visual outputs, and model performance rather than code-only execution. Our extensive experiments with 15 closed- and open-source models show that even the strongest agent, Claude-4.6-Sonnet, achieves only 56.70% task success, while all open-source agents remain below 1%, frequently failing at tool orchestration, OS grounding, and multi-step reasoning. These results reveal a substantial capability gap between current agentic systems and real data-science workflows, positioning DSAgentBench as a foundation for developing grounded, verifiable, autonomous data-science agents. We release DSAgentBench at https://github.com/vis-nlp/DSAgentBench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑