表格基础模型仍然需要特征工程吗?
Do Tabular Foundation Models Still Need Feature Engineering?
浏览论文内容
中文总结 AI 辅助
本研究通过受控实验发现,随着表格基础模型能力增强,特征工程带来的性能提升逐渐消失,而补充相关数据集上下文仍能改善性能,表明性能增益来源已从输入重表示转向任务相关上下文。
中文摘要 AI 辅助
特征工程长期以来一直是表格机器学习的一个基石。表格基础模型(TFMs)在广泛的表格数据集上进行预训练,并通过上下文学习进行应用。它们的兴起提出了一个自然的问题:随着这些模型变得越来越强大,手动特征构建是否仍然重要?为了回答这个问题,我们在两个主要TFM家族的多个版本上进行了受控研究,在TabArena的基准数据集上测试了多种现有的特征工程技术。我们发现了一个一致的规律:特征工程的收益集中在较早的模型世代中,而对于最强的模型来说,这些收益变得可以忽略不计。这些结果表明,更强的TFMs对显式工程化的输入表示的依赖较少。然而,在一项补充实验中,从相关数据集添加上下文信息仍然可以提高性能。我们的发现表明,对于更强的TFMs,性能提升的来源发生了转变:重新表示现有输入变得不那么有效,而提供额外的任务相关上下文仍然是有益的。
英文摘要
Feature engineering has long been a cornerstone of tabular machine learning. Tabular foundation models (TFMs) are pretrained on a wide range of tabular datasets and applied via in-context learning. Their rise raises a natural question: does manual feature construction still matter as these models become more capable? To answer this, we perform a controlled study across several versions of two major TFM families, testing a wide range of existing feature engineering techniques on benchmark datasets from TabArena. We find a consistent pattern: feature engineering gains are concentrated in earlier model generations and become negligible for the strongest models. These results suggest that stronger TFMs depend less on explicitly engineered input representations. In a complementary experiment, however, adding in-context information from related datasets still improves performance. Our findings indicate a shift in the source of performance gains for stronger TFMs: re-representing existing inputs becomes less effective, while providing additional task-relevant context remains beneficial.