表格基础模型的出人意料的泛化特性研究
Understanding the Surprising Generalization Properties of Tabular Foundation Models
- Polytechnique Montréal(蒙特利尔理工学院)
- Mila – Quebec AI Institute(米拉-魁北克人工智能研究所)
- Chandar Research Lab(钱达尔研究实验室)
- University of Toronto(多伦多大学)
- Layer 6 AI(第六层人工智能公司)
- Prior Labs(普里奥实验室)
- ELLIS Institute Tübingen(埃利斯研究所图宾根分部)
- University of Freiburg(弗赖堡大学)
- Cohere(科here公司)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究揭示仅在单个真实表格上预训练的TFMs有强迁移性,提出以任务为中心、基于检索的新视角,为TFMs的模型与语料库设计提供了新框架。
中文摘要 AI 辅助
表格基础模型(Tabular Foundation Models, TFMs)越来越多地依赖上下文学习,即在推理时模型接收带标签示例,无需更新权重即可为新输入预测标签。现有TFMs通常在大量合成语料库或超大型真实数据集集合上训练,与之不同,本文表明仅在单个真实表格上进行自监督预训练就能产生出人意料的强迁移效果。在该设置下,我们还发现无论下游预测任务如何,表格往往要么广泛有用要么广泛无用,且有用性的最强预测因子是特征数量而非实例数量。这引出了表格预训练的以任务为中心的解释:任务的数量和质量对TFMs的预训练至关重要。我们表明,这种以任务为中心的视角可助力大规模语料库设计:细粒度的列级预处理始终能提升下游性能,而在数据集层面进行过滤或去重时则未观察到性能提升。最后,我们为TFMs的泛化提供了新视角:我们认为表格上下文泛化在很大程度上是基于检索的,良好的模型是那些能学会识别提供的上下文内相关示例并对其进行良好聚合的模型。TFMs的机制相对未被充分研究;我们以任务为中心、基于检索的视角为未来的模型和语料库设计提供了新框架。
英文摘要
Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs without updating its weights. Existing TFMs are typically trained on either massive synthetic corpora or very large collections of real datasets. In contrast, we show that surprisingly strong transfer can emerge from self-supervised pre-training on just a single real table. In this setting, we also find that tables tend to be either broadly useful or broadly poor regardless of downstream prediction task, and that the strongest predictor of usefulness is the number of features rather than the number of instances. This leads to a task-centric interpretation of tabular pre-training: the number and the quality of tasks are essential for the pre-training of TFMs. We show that the same task-centric perspective can help corpus design at scale: fine-grained column-level pre-processing consistently improves downstream performance, while no improvements are observed when we filter or deduplicate at the dataset level. Finally, we offer a new perspective for how TFMs generalize: we believe that tabular in-context generalization is largely retrieval-based, and good models are those that learn to identify relevant examples in the provided context and aggregate them well. The mechanics of TFMs have been relatively understudied; our task-centric, retrieval-based perspective offers a new framework to guide future model and corpus design.