DS-Lighting:让智能体管控框架显式化以实现数据科学自动化
DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
浏览论文内容
中文总结 AI 辅助
DS-Lighting是一款显式化数据科学自动化管控框架的工具包,将管控框架分解为四层,集成多基准实现可控比较,可提升结果的可复现性、可比性与可靠性,减少系统级故障。
中文摘要 AI 辅助
大型语言模型(LLM)智能体在自动化数据科学工作流方面展现出潜力,但其端到端性能关键依赖于代表任务、管理执行状态、约束输出工件并提供评估反馈的智能体管控框架。现有数据科学智能体常将该管控框架隐式化,导致跨异构任务的结果难以复现、比较与归因。我们提出DS-Lighting,这是一款为数据科学自动化使管控框架设计显式化的统一管控框架工具包。DS-Lighting将管控框架分解为四个可复用层:数据层、工作流层、执行层与评估层,并将各类智能体表示为可执行算子程序,支持预定义管道与自适应搜索。我们进一步将多个开源数据科学基准集成至MLE-Bench风格的任务格式,实现在共享任务接口、沙盒运行时与度量协议下的可控比较。跨智能体、管控框架、模型及消融实验的结果表明,显式管控框架设计可提升可复现性、可比性与可靠性,同时减少端到端数据科学工作流中可避免的系统级故障。我们的代码可在该httpsURL获取。
英文摘要
Large Language Model (LLM) agents have shown promise for automating data-science workflows, yet their end-to-end performance depends critically on the agent harness that represents tasks, manages execution state, constrains output artifacts, and provides evaluation feedback. Existing data-science agents often leave this harness implicit, making results difficult to reproduce, compare, and attribute across heterogeneous tasks. We introduce DS-Lighting, a unified harness toolkit that makes harness design explicit for data-science automation. DS-Lighting decomposes the harness into four reusable layers: data, workflow, execution, and evaluation, and represents diverse agents as executable operator programs that support both predefined pipelines and adaptive search. We further integrate multiple open-source data-science benchmarks into an MLE-Bench-style task format, enabling controlled comparison under a shared task interface, sandboxed runtime, and metric protocol. Experiments across agents, harnesses, models, and ablations show that explicit harness design improves reproducibility, comparability, and reliability, while reducing avoidable system-level failures in end-to-end data-science workflows. Our code is available at https://github.com/usail-hkust/dslighting
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。