arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自监督预训练何时对表格模型有帮助?标签稀缺与缺失数据的研究

When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data

Sahand Mazrouei

arXiv 2608.24381首次发表:更新:

发表机构

Faculty of Mathematics and Computer Science, Kharazmi University(卡拉兹米大学数学与计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究了标签稀缺与缺失数据场景下表格SSL预训练的效果,发现其在干净数据集上改进可靠,在高缺失数据集上常失效,且掩码并恢复目标与其他表格SSL基线无显著差异。

AI 中文摘要

自监督学习(SSL)已成为表格数据领域颇具前景的方法,但在极端标签稀缺和测试时数据缺失场景下的有效性仍未得到充分探索。本文在14项不同的分类任务中,对比了一种“掩码并恢复”的SSL预训练目标、从头训练以及经典基线方法的性能。首先,SSL在平均表现上优于从头训练,且与最先进的树集成模型具有竞争力(在10%标签设置下,SSL的AUC约为0.8954,随机森林的AUC为0.9015),但SSL与从头训练的性能增益存在较高的任务间方差,且不具备统计显著性(在5%和10%标签设置下p值均为0.626)。其次,与“缺失值补全目标普遍对存在原生缺失的数据集有益”的假设相反,SSL在干净数据集上的改进最为可靠,而在固有缺失率高的数据集上常导致性能下降。第三,尽管存在上述训练方差,SSL预训练模型在两种测试时缺失场景下的平均AUC均高于从头训练模型:完全随机缺失(MCAR)注入时,平均AUC提升0.0245,在14项任务中11项表现更优;结构化缺失偏移(MNAR)时,平均AUC提升0.0418,在14项任务中8项表现更优,但经Holm-Bonferroni多重比较校正后,两种场景下的差异均不具备统计显著性(校正后p值分别为0.118和0.518)。第四,在相同编码器架构下,将本文的“掩码并恢复”目标与三种已有的表格SSL基线方法(VIME、SCARF、SubTab)对比,发现与它们均无显著差异(校正后p值分别为0.459、1.000、1.000),表明本文的发现反映了表格SSL的通用特性,而非某一特定预文本任务的特有属性。

英文摘要

Self-supervised learning (SSL) has emerged as a promising approach for tabular data, yet its efficacy under extreme label scarcity and test-time missingness remains under-explored. In this paper, we evaluate a mask-and-recover SSL pretraining objective against training from scratch and classical baselines across 14 diverse classification tasks. First, while SSL outperforms training from scratch on average and remains competitive with state-of-the-art tree ensembles (achieving ~0.8954 AUC vs. Random Forest's 0.9015 at 10% labels), the SSL-vs-scratch gains exhibit high inter-task variance and lack significance (p = 0.626 at both 5% and 10% labels). Second, contrary to the hypothesis that missing-value imputation objectives universally benefit datasets with native missingness, SSL yields the most reliable improvements on clean datasets, while frequently degrading performance on datasets with high inherent missingness. Third, despite this training variance, SSL-pretrained models achieve a higher average AUC than scratch-trained models under both test-time missingness completely at random (MCAR) injection (+0.0245 AUC, positive on 11 of 14 tasks) and structured missingness shifts (MNAR, +0.0418 AUC, positive on 8 of 14 tasks), though neither difference remains statistically significant after Holm-Bonferroni correction for multiple comparisons (adjusted p = 0.118 and p = 0.518, respectively). Fourth, comparing our mask-and-recover objective against three established tabular SSL baselines (VIME, SCARF, SubTab) under an identical encoder architecture, we find no significant difference from any of them (adjusted p = 0.459, p = 1.000, p = 1.000), indicating our findings reflect general properties of tabular SSL rather than idiosyncrasies of one particular pretext task.

Comments18 pages, 4 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑