arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21894cs.LG

LLMs作为文本与表格预测的特征工程师

LLMs as Feature Engineers for Text-and-Tabular Prediction

Merwan Barlier, Blaz Skrlj

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一个迭代框架,利用LLM自动从文本中提取可解释特征用于表格预测,通过错误反馈加速特征发现,并在多数据集上提升性能与可解释性。

中文摘要 AI 辅助

我们引入了一个迭代框架,该框架自动化地从非结构化文本中提取可解释的、受模式约束的分类特征,用于表格预测模型。为了导航特征空间,一个生成器LLM提出语义定义,一个独立的提取器LLM实现特征,下游表格模型评估其预测性能。我们通过将显式模型错误(如AUC排名反转)转化为自然语言反馈来优化这一搜索,引导LLM解决特定的预测失败。在三个公开数据集上的评估表明,与无引导搜索相比,这种错误驱动的循环将特征发现速度提高了最多3倍。实验上,生成的特征表现出强大的多视图互补性,在与TF-IDF和稠密嵌入结合时严格优于任何子集。最后,该框架保证了实例级别的可解释性:发现的特征在SHAP重要性排名中占主导地位,并为每个预测提供了完全透明、语义化的审计轨迹。

英文摘要

We introduce an iterative framework that automates the extraction of interpretable, schema-bound categorical features from unstructured text for tabular prediction models. To navigate the feature space, a generator LLM proposes semantic definitions, a separate extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance. We optimize this search by translating explicit model errors, such as AUC ranking inversions, into natural-language feedback, steering the LLM to resolve specific predictive failures. Evaluated across three public datasets, this error-driven loop accelerates feature discovery by up to $3\times$ compared to unguided search. Empirically, the generated features demonstrate strong multi-view complementarity, strictly outperforming any subset when combined with TF-IDF and dense embeddings. Finally, the framework guarantees instance-level interpretability: the discovered features dominate SHAP importance rankings and provide a fully transparent, semantic audit trail for every prediction.

发表机构

  • Teads

机构由 AI 辅助整理,请以论文原文为准。

↑