TabFM:面向表格数据的零样本基础模型
TabFM: A Zero-Shot Foundation Model for Tabular Data
浏览论文内容
中文总结 AI 辅助
TabFM是一个400M参数的表格基础模型,通过上下文学习实现零样本预测,在51个基准数据集上超越AutoML,并提出了两个扩展进一步提升性能。
中文摘要 AI 辅助
表格机器学习通常依赖于每个数据集的工作流程,针对每个任务从头拟合树集成或运行AutoML搜索。我们提出了TabFM,一个400M参数的表格基础模型,将监督式表格预测形式化为上下文学习。TabFM在单次前向传播中产生校准的零样本预测,无需针对特定任务进行调优。TabFM完全在由结构因果模型生成的合成表格上训练,学习通用的表格表示,这些表示可以零样本迁移到现实世界任务。在TabArena中所有51个基准数据集(38个分类和13个回归)上,零样本TabFM在默认表格基础模型中排名第一,并优于调优后的AutoML流水线。基于相同冻结权重的两个扩展在两个轨道上进一步提升了性能:多视图特征扩展与集成及事后校准(TabFM+),以及LLM引导的、数据集特定的数据处理和特征工程(TabFM-Auto)。
英文摘要
Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representations that transfer zero-shot to real-world tasks. Across all 51 benchmark datasets in TabArena (38 classification and 13 regression), zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines. Two extensions over the same frozen weights improve performance further on both tracks: multi-view feature expansion with ensembling and post-hoc calibration (TabFM+), and LLM-guided, dataset-specific data processing and feature engineering (TabFM-Auto).
发表机构
- Google Research(谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。