arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用大语言模型进行特征生成:一种进化算法方法

Feature Generation Using LLMs: An Evolutionary Algorithm Approach

Aria Nourbakhsh, Benoît Alcaraz, Christoph Schommer

arXiv 2607.16255首次发表:更新:

发表机构

University of Luxembourg(卢森堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究利用大语言模型解决表格数据特征生成问题,创建管道结合属性与提示生成新特征,经选择算法筛选,应用于八个不同数据集,结果显示语言模型生成的新特征有助于给定任务并提升分类效果。

AI 中文摘要

机器学习管道中的关键步骤是为每个实体提供能代表已处理实体特征的特征或属性。特征工程是找出属性间关系的重要步骤,否则这些属性可能无法被机器学习算法处理。同时,大语言模型在编码、数学推理和处理世界知识方面展现出了有前景的能力。在这项工作中,我们基于先前给定的特征,利用大语言模型解决表格数据的特征生成问题。我们创建了一个管道,它接收一组属性和一个提示来生成新特征。然后,我们的选择算法选择性能最佳的属性集。我们将方法应用于八个来自不同领域和数据类型的数据集。结果表明,在大多数情况下,语言模型能基于数学和逻辑运算符生成对给定任务有用的新特征,并能改善分类结果。

英文摘要

A crucial step in machine learning pipelines is to present each entity with features or attributes that are representative of the characteristics of the processed entities. Feature engineering is an important step in finding a relation among attributes that otherwise may not be processed by the ML algorithms. Meanwhile, Large Language Models have shown promising abilities in coding, mathematical reasoning, and processing world knowledge. In this work, we utilize an LLM for the problem of feature generation from tabular data based on the previously given features. We have created a pipeline that takes a set of attributes and a prompt to generate new features. Then, our selection algorithm selects the best-performing sets of attributes. We apply our method to eight datasets from different domains and data types. Our results show that, in most cases, the language model can produce new features based on mathematical and logical operators that are useful for the given tasks and can improve classification results.

Journal refCommun. Comput. Inf. Sci. 2471, 48-64 (2025)

DOI:10.1007/978-3-031-89103-8_4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑