arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DataFlow-Harness:用于构建可编辑大语言模型数据管道的基础代码代理平台

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang

arXiv 2607.16617首次发表:更新:

发表机构

Peking University; Institute for Advanced Algorithms Research; Zhongguancun Academy(北京大学; 先进算法研究所; 中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型数据处理工作流中编码代理脚本无法自动转化为可编辑工件的问题,提出DataFlow-Harness平台,通过类型化增量突变引导构建DAG,结合多种技术,在数据工程基准测试中取得高通过率,降低成本和延迟。

AI 中文摘要

大语言模型(LLMs)越来越多地用于自动化数据处理工作流程,但编码代理通常生成的脚本不会自动作为持久、可编辑的平台工件实现。我们称这种脱节为“NL2Pipeline差距”。为了弥合这一差距,我们引入了DataFlow-Harness平台,该平台通过类型化、增量式突变引导大语言模型代理构建平台原生有向无环图(DAG),而非自由形式脚本。该平台结合了用于过程指导的DataFlow-Skills、暴露实时运算符注册表和当前管道状态的模型上下文协议(MCP)层,以及将对话式创作与可视化DAG编辑器同步的DataFlow-WebUI。在12任务数据工程基准测试中,DataFlow-Harness实现了93.3%的观察到的端到端通过率。相对于Vanilla Claude Code,它将测量的货币成本降低了72.5%,生成延迟降低了49.9%;其观察到的通过率与上下文感知Claude Code基线相差0.9个百分点以内,而成本低42.8%。每项任务分析表明,当构建依赖于隐式过程知识时,技能最有用。这些结果表明,实时平台基础可以产生持久、可编辑的工作流工件,观察到的可靠性接近脚本生成基线,且测量的构建成本和延迟更低。

英文摘要

Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the \textit{NL2Pipeline gap}. To bridge it, we introduce \textsc{DataFlow-Harness}, a platform that guides an LLM agent to construct platform-native directed acyclic graphs (DAGs) through typed, incremental mutations rather than free-form scripts. The platform combines \textsc{DataFlow-Skills} for procedural guidance, a Model Context Protocol (MCP) layer that exposes the live operator registry and current pipeline state, and \textsc{DataFlow-WebUI}, which synchronizes conversational authoring with a visual DAG editor. On a 12-task data-engineering benchmark, \textsc{DataFlow-Harness} achieves a 93.3\% observed end-to-end pass rate. Relative to Vanilla Claude Code, it reduces measured monetary cost by 72.5\% and generation latency by 49.9\%; its observed pass rate is within 0.9 percentage points of the Context-Aware Claude Code baseline while its cost is 42.8\% lower. Per-task analysis indicates that Skills are most useful when construction depends on implicit procedural knowledge. These results show that live platform grounding can produce persistent, editable workflow artifacts with an observed reliability close to script-generation baselines and with lower measured construction cost and latency.

Comments13 pages, 2 figures, and 5 tables. Technical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑