arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在小型表格数据集上理解TabPFN的上下文采样

Understanding Context Sampling in TabPFN on Small Tabular Datasets

Mohammed Abdullah

arXiv 2607.26628首次发表:更新:

发表机构

Anna University(安娜大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对小型表格数据集,探究TabPFN上下文采样的上下文大小、选择方式对预测性能的影响,发现更大上下文更稳定准确,随机采样因覆盖特征空间更优,高成本选择方法无优势。

AI 中文摘要

TabPFN通过上下文学习执行分类:它以一组带标签的训练行(即上下文或原型)为条件,无需梯度更新即可预测测试标签。在小型表格数据集上,从业者仍需选择上下文大小以及哪些行构成上下文。我们通过在15个OpenML数据集上重复进行上下文采样,研究这些选择如何影响预测稳定性、准确率和选择成本。具体而言,我们调查了三个问题:(i)更大的上下文是否会减少随机抽取带来的预测变异性;(ii)准确率是否取决于训练分布的保留还是特征空间的覆盖;(iii)K-Means和最远点采样等成本较高的选择方法是否比均匀随机采样更具优势。我们发现,更大的上下文既更准确也显著更稳定,在有改进空间的数据集上,当k=16时AUC变异系数约为6-18%,在更大的上下文大小下降至1-4%。尽管准确率与随机上下文的分布代表性相关,但对照实验显示,仅匹配特征均值会使AUC降低多达0.5,因为这会减少上下文多样性。混合效应分析确定,多样性和覆盖度而非特征均值匹配是准确率的更强预测因子(多样性β=+0.23,p=3×10^-12;特征均值偏移β=-0.01,p=0.71)。K-Means和最远点采样的准确率与随机选择相似,但选择成本高两个到三个数量级。这些结果表明,随机采样之所以成功,是因为它有望提供特征空间覆盖,而非因为它能再现潜在数据分布。

英文摘要

TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. On small tabular datasets, practitioners must still choose the context size and which rows constitute the context. We study how these choices affect prediction stability, accuracy, and selection cost using repeated context sampling on 15 OpenML datasets. Specifically, we investigate (i) whether larger contexts reduce prediction variability across random draws, (ii) whether accuracy depends on preserving the training distribution or on feature-space coverage, and (iii) whether expensive selection methods such as K-Means and farthest-point sampling provide benefits over uniform random sampling. We find that larger contexts are both more accurate and substantially more stable, with AUC coefficient of variation decreasing from roughly 6 to 18% at k=16 to 1 to 4% at larger context sizes on datasets with room for improvement. Although accuracy correlates with distribution representativeness in random contexts, controlled experiments show that matching feature means alone can reduce accuracy by up to 0.5 AUC because it reduces context diversity. Mixed-effects analysis identifies diversity and coverage, rather than feature-mean matching, as the stronger predictor of accuracy (diversity beta=+0.23, p=3x10^-12; feature-mean shift beta=-0.01, p=0.71). K-Means and farthest-point sampling achieve similar accuracy to random selection while requiring two to three orders of magnitude more selection cost. These results show that random sampling succeeds because it provides feature-space coverage in expectation, not because it reproduces the underlying data distribution.

Comments12 pages, 4 figures. Code and experiment logs available at https://github.com/mohammed1916/tabmx

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑