arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19754cs.AIcs.CL

AutoData:用于预训练数据选择的智能体搜索

AutoData: Agentic Search for Pre-training Data Selection

  • University of Amsterdam(阿姆斯特丹大学)
  • Weco AI

机构由 AI 辅助整理,请以论文原文为准。

Yan Meng, Dhruv Srikanth, Bingchen Zhao, Zhengyao Jiang, Yuxiang Wu

中文总结 AI 辅助

AutoData将预训练数据选择视为智能体搜索问题,通过迭代优化可执行选择算法,发现优于人工流程的配方并提升下游指标。

中文摘要 AI 辅助

近年来,LLM智能体通过在执行反馈下编辑模型和训练代码,在自动化机器学习工程方面展现出前景。然而,数据在很大程度上仍处于这种智能体优化循环之外。我们将预训练数据选择定义为对每文档特征(即词汇统计、分类标签和困惑度)的启发式工程。我们引入了AutoData,一个直接在可执行选择算法上进行搜索的智能体。与先前在固定领域集合上优化权重的数据混合方法不同,AutoData搜索一个更丰富的程序空间,包括评分、分层和随机选择规则,通过使用代理模型的验证反馈迭代改进算法,自动发现特征交互。在一次过夜搜索中,AutoData发现了一种选择算法,其性能优于现有的人工设计的数据筛选流程。尽管仅在此小型代理上搜索,所发现的配方可迁移到更大规模,并提高了下游指标CORE。这些结果表明,数据工程可以被视为一个智能体机器学习问题,将自主研究从模型和训练代码优化扩展到数据领域。

英文摘要

LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.

↑