面向非表格任务的表格基础模型
Tabular foundation models for non-tabular tasks
浏览论文内容
中文总结 AI 辅助
该研究探究表格基础模型TabPFN v3是否可用于非表格任务,将MNIST手写数字识别、法德语言识别、Tiny ImageNet图像分类转为表格形式,其在部分任务上达到了对应专用模型的相当准确率。
中文摘要 AI 辅助
表格基础模型(Tabular foundation models, TFMs)是近期出现的一种针对表格数据机器学习的有前景范式,具备无需针对特定任务训练即可跨数据集泛化的能力。由于许多机器学习数据集可表示为表格,这引发了一个问题:TFM的能力是否能延伸至传统上被视为非表格的任务?我们通过在三个非表格分类问题上使用TabPFN v3来解决该问题:MNIST手写数字识别、法语与德语单词的语言识别,以及Tiny ImageNet图像分类。在每种情况下,原始数据都被表示为表格的行,分类被表述为对缺失标签的预测。我们评估了预训练模型在提供不同数量上下文样本时的性能,未进行额外训练或微调。尽管未明确访问表征数据的空间或序列结构,TabPFN v3在部分案例中达到了与针对对应任务设计的模型或方法相当的准确率。
英文摘要
Tabular foundation models (TFMs) have recently emerged as a promising paradigm for machine learning on tabular data, offering the ability to generalize across datasets without task-specific training. Since many machine learning datasets can be represented as tables, this raises the question: does TFM capability extend beyond tasks traditionally regarded as tabular? We address this question by using TabPFN v3 on three non-tabular classification problems: handwritten digit recognition on MNIST, language identification of French and German words, and image classification on Tiny ImageNet. In each case, the original data are represented as rows of a table and classification is formulated as prediction of a missing label. We evaluate performance as a function of the number of context samples provided to the pretrained model, with no additional training or fine-tuning. Despite having no explicit access to the spatial or sequential structure characterizing the data, TabPFN v3 in some cases achieves accuracies comparable with that of models or methods geared specifically toward the corresponding tasks.
发表机构
- Technische Universität Dresden(德累斯顿工业大学)
- Maynooth University(梅努斯大学)
- Universität Würzburg(维尔茨堡大学)
机构由 AI 辅助整理,请以论文原文为准。