arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28408cs.LG

SymboLLM-FE:用于表格数据自动特征工程的大语言模型加速符号回归

SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data

Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo

首次发表
浏览论文内容

中文总结 AI 辅助

SymboLLM-FE将符号回归与LLM结合,通过符号回归提取强相关公式并经LLM优化,在6个真实数据集和4个Kaggle竞赛上性能优于现有AutoFE,解决了可解释性差和迭代多的问题。

中文摘要 AI 辅助

表格数据是机器学习的核心数据格式,但由于特征信息不足,往往缺乏高性能建模所需的判别能力。自动特征工程(AutoFE)通过自动化特征生成与选择解决这一问题,兼顾模型性能与运算效率。然而,传统AutoFE依赖盲目数学变换,生成的特征可解释性差;基于大语言模型(LLM)的AutoFE则存在需多轮高成本迭代才能生成高效用特征以提升模型性能的挑战,还存在固有偏差与幻觉风险。本文将符号回归与LLM结合用于特征工程(SymboLLM-FE)以解决上述挑战:通过符号回归提取与目标强相关的具有数学表达力的公式,可提升模型性能,再利用具备丰富先验知识的LLM对这些公式进行优化,确保可解释性。在6个真实世界数据集和4个Kaggle竞赛上的实验结果表明,SymboLLM-FE的性能优于现有AutoFE方法,且通过采用基于统计先验的LLM优化机制和仅需个位数的LLM调用,解决了可解释性差和迭代次数多的双重挑战。

英文摘要

Tabular data, as a core data format in machine learning, often lacks the discriminative power needed for high-performance modeling due to insufficient feature informativeness. Automated Feature Engineering (AutoFE) overcomes this by automating feature generation and selection, ensuring both model performance and operational efficiency. However, traditional AutoFE often yield features with poor interpretability because they rely on blind mathematical transformations, while large language models (LLM)-based AutoFE faces challenges in requiring costly multi-round iterations to generate high-utility features to effectively enhance model performance, compounded by inherent risks of bias and hallucination. In this paper, we combine symbolic regression with LLMs for feature engineering (SymboLLM-FE) to solve these challenges. We extract mathematically expressive formulas strongly correlated with the target via symbolic regression, which can enhance model performance, then refine them by LLMs with rich prior knowledge to ensure interpretability. Empirical results on six real-world datasets and four Kaggle competitions demonstrate that SymboLLM-FE outperforms existing AutoFE. SymboLLM-FE also addresses the dual challenges of poor interpretability and numerous iterations by employing a statistical prior-grounded LLM refinement mechanism and single-digit LLM calls.

发表机构

  • Nanjing University(南京大学)
  • School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
  • School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑