arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37959cs.LG

TabFM:面向表格数据的零样本基础模型

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

首次发表
浏览论文内容

中文总结 AI 辅助

TabFM是一个400M参数的表格基础模型,通过上下文学习实现零样本预测,在51个基准数据集上超越AutoML,并提出了两个扩展进一步提升性能。

中文摘要 AI 辅助

表格机器学习通常依赖于每个数据集的工作流程,针对每个任务从头拟合树集成或运行AutoML搜索。我们提出了TabFM,一个400M参数的表格基础模型,将监督式表格预测形式化为上下文学习。TabFM在单次前向传播中产生校准的零样本预测,无需针对特定任务进行调优。TabFM完全在由结构因果模型生成的合成表格上训练,学习通用的表格表示,这些表示可以零样本迁移到现实世界任务。在TabArena中所有51个基准数据集(38个分类和13个回归)上,零样本TabFM在默认表格基础模型中排名第一,并优于调优后的AutoML流水线。基于相同冻结权重的两个扩展在两个轨道上进一步提升了性能:多视图特征扩展与集成及事后校准(TabFM+),以及LLM引导的、数据集特定的数据处理和特征工程(TabFM-Auto)。

英文摘要

Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representations that transfer zero-shot to real-world tasks. Across all 51 benchmark datasets in TabArena (38 classification and 13 regression), zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines. Two extensions over the same frozen weights improve performance further on both tracks: multi-view feature expansion with ensembling and post-hoc calibration (TabFM+), and LLM-guided, dataset-specific data processing and feature engineering (TabFM-Auto).

发表机构

  • Google Research(谷歌研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑