arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29966cs.CL

DataFoundry:通过递归自改进演化数据准备器

DataFoundry: Evolving Data Preparators via Recursive Self-Improvement

Cehao Yang, Xiaojun Wu, Xueyuan Lin, Chengjin Xu, Xuhui Jiang, Hui Xiong, Jian Guo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出DataFoundry框架,通过Skills-as-Modules架构的递归自改进演化数据准备器,在DataPrep-Bench多领域基准上验证其生成训练数据的下游效用优于基线,且改进不依赖特定模型。

中文摘要 AI 辅助

大型语言模型的领域适应越来越依赖于构建高质量的训练数据,但现有的数据准备管道通常仅在生成后通过事后过滤来处理质量问题。这造成了根本的不匹配:数据质量问题往往源于构建过程本身,而质量控制仅应用于其输出。我们引入DataFoundry,这是一个在大规模数据生产前通过递归自改进演化数据准备器的框架。DataFoundry将数据准备器表示为可演化的运行时规范,并通过Skills-as-Modules架构实例化其演化,其中中央Controller协调模块化技能以编译可执行运行时,使用领域适用的标准在小型试点集上诊断缺陷,并将诊断反馈转换为适配器,该适配器在保留稳定接口的同时修改各个准备组件。我们在涵盖数学、金融、法律和医学的DataPrep-Bench上评估DataFoundry,发现递归演化的准备器生成的训练数据比基线具有更高的下游效用。跨不同主干的实验进一步表明,这些改进不依赖于特定模型,而分析和案例研究进一步揭示了该框架的优化动态,并说明了其演化在实践中如何展开。

英文摘要

Domain adaptation of large language models increasingly depends on constructing high-quality training data, yet existing data-preparation pipelines typically address quality only after generation through post-hoc filtering. This creates a fundamental mismatch: data-quality issues often originate from the construction process itself, while quality control is applied only to its outputs. We introduce \textsc{DataFoundry}, a framework for \textbf{evolving data preparators through recursive self-improvement} before large-scale data production. \textsc{DataFoundry} represents a data preparator as an evolvable runtime specification and instantiates its evolution with a \textsc{Skills-as-Modules} architecture, in which a central \textsc{Controller} orchestrates modular skills to compile executable runtimes, diagnose deficiencies on small pilot sets using domain-appropriate criteria, and translate diagnostic feedback into adapters that revise individual preparation components while preserving stable interfaces. We evaluate \textsc{DataFoundry} on DataPrep-Bench across mathematics, finance, law, and medicine, and find that recursively evolved preparators produce training data with higher downstream utility than baselines. Experiments across different backbones further demonstrate that these improvements are not tied to a particular model, while analyses and case studies further reveal the framework's optimization dynamics and illustrate how its evolution unfolds in practice.

发表机构

  • IDEA Research, International Digital Economy Academy(国际数字经济研究院IDEA研究院)
  • Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • DataArc Tech Ltd.(DataArc科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑