arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16416cs.LGcs.NE

用语言模型演化用于自动机器学习的可执行流水线程序

Evolving Executable Pipeline Programs for AutoML with Language Models

Sofoklis Kitharidis, Cor J. Veenman, Jan N. van Rijn, Thomas Bäck, Niki van Stein

首次发表
浏览论文内容

中文总结 AI 辅助

LACE是首个以代码形式构建通用表格型AutoML的框架,用语言模型演化可执行流水线,在68个OpenML分类任务上表现优异,其核心贡献是提供可直接复用的代码级搜索空间。

中文摘要 AI 辅助

自动机器学习(AutoML)系统会在预先指定的预处理算子、学习器和超参数空间内搜索流水线,它们可以选择并调优已知组件,但无法生成该空间之外的结构。我们提出LACE,一种AutoML框架,它转而在完整的可执行流水线程序上进行搜索:演化循环维护着一组兼容scikit-learn的Python类,而大型语言模型充当变异算子。据我们所知,LACE是首个以这种方式构建通用表格型流水线AutoML的框架,在标准化OpenML任务上,采用了向生成器隐藏数据集身份的泄漏控制协议进行评估。由于每个候选都是普通Python代码,返回的流水线及其生成过程可被直接检查和编辑,而非仅通过框架的模型对象。在68个OpenML分类任务上,搭载GPT-5.4-mini的LACE显著优于auto-sklearn、H2O和固定XGBoost基线,与评估中最强的基于搜索的系统AutoGluon相比无显著差异,且覆盖了全部基准。更新的表格型基础模型在其支持的任务子集上更准确,但它们应用固定预训练预测器,而非返回可编辑的特定任务程序。因此,LACE的贡献不在于原始准确率,而在于由代码定义的搜索空间:完整覆盖、从业者可直接复用的流水线,以及通过编辑提示而非框架来扩展的组件集。

英文摘要

Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, learners, and hyper-parameters specified in advance: they can select and tune known components, but cannot produce structure outside that space. We present LACE, an AutoML framework that instead searches over complete executable pipeline programs: an evolutionary loop maintains a population of scikit-learn-compatible Python classes, and a large language model acts as the variation operator. To our knowledge, LACE is the first to formulate general tabular pipeline AutoML this way, evaluated on standardized OpenML tasks under a leakage-controlled protocol that withholds dataset identity from the generator. Because every candidate is ordinary Python, the returned pipeline and the search that produced it can be inspected and edited directly, rather than only through a framework's model objects. On 68 OpenML classification tasks, LACE with GPT-5.4-mini significantly outperforms auto-sklearn, H2O, and a fixed XGBoost baseline, with no detectable difference against AutoGluon, the strongest search-based system evaluated, while covering the full benchmark. Newer tabular foundation models are more accurate on the subset of tasks they support, but apply a fixed pretrained predictor rather than returning an editable task-specific program. LACE's contribution is therefore not raw accuracy but a search space defined by code: complete coverage, pipelines practitioners can reuse directly, and a component set extended by editing the prompt rather than the framework.

↑