arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03500cs.LGstat.ML

深度表格生成器在何种训练规模下不再优于平凡基线?——基于临床与标准数据集规模阶梯的预注册基准测试

Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets

Shivam Shrivastava

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过预注册的规模阶梯基准测试发现,深度表格生成器在训练样本少于20,000行时通常无法超越平凡基线,且排名稳定性在小规模下反而更高,质疑了子采样大数据集作为小数据集替代品的有效性。

中文摘要 AI 辅助

深度表格生成模型通常在拥有数万行数据的数据集上进行基准测试;而临床数据集通常只有数百行。我们预注册并运行了一个规模阶梯基准测试,以找出两种场景的分界点:8个公共数据集从200行到20,000行训练样本进行子采样,7种生成器(独立边际分布、高斯copula、SMOTE、非条件SMOTE、CTGAN、TVAE、TabDDPM)使用固定的20次调参预算和5个评估种子,外加4个真实规模的原生小临床数据集,总计2,220次承诺运行。主要指标是固定分类器在合成数据上训练并在真实数据上测试的AUROC。在24个(数据集,深度模型)配对中,有23个在任何我们测量的训练规模下,深度模型从未超过最佳平凡基线超过种子噪声。最佳基线在49个(数据集,规模)单元中赢得40个。我们预注册的预测——深度模型在较小规模下的排名不稳定——被证伪:在5,000行以下的相邻梯级之间,平均Kendall tau为0.806,高于我们0.8的阈值,且稳定性在最小规模处最高而非最低。一个注意事项限制了所有这些结果:在81%的单元中,前两名方法之间的差距小于种子之间的差异。最后,在原生小临床数据集上的方法排名与在子采样大数据集上的排名仅中等程度一致(平均tau为0.57至0.64),这质疑了子采样大数据集能否代表小数据集。所有2,220个结果文件、预注册及其哈希值,以及从这些文件重新生成每个图表和数字的代码均已公开。

英文摘要

Deep tabular generative models are benchmarked on datasets with tens of thousands of rows; clinical datasets have hundreds. We preregistered and ran a size-ladder benchmark to find where the two regimes diverge: 8 public datasets subsampled from 200 to 20,000 training rows, seven generators (independent marginals, Gaussian copula, SMOTE, unconditional SMOTE, CTGAN, TVAE, TabDDPM) with a fixed 20-trial tuning budget and 5 evaluation seeds, plus 4 natively small clinical datasets at true size, for 2,220 committed runs in total. The primary metric is the AUROC of fixed classifiers trained on synthetic and tested on real data. In 23 of 24 (dataset, deep model) pairs no deep model ever beats the best trivial baseline by more than seed noise, at any training size we measured. The best baseline wins 40 of 49 (dataset, size) cells. Our preregistered prediction that the deep models' ranking would be unstable at small sizes is falsified: mean Kendall tau between adjacent rungs below 5,000 rows is 0.806, above our 0.8 threshold, and stability is highest at the smallest sizes rather than lowest. One caveat bounds all of this: in 81% of cells the gap between the top two methods is smaller than the variation between seeds. Finally, method rankings on natively small clinical datasets agree only moderately with rankings on subsampled large ones (mean tau 0.57 to 0.64), which questions whether a subsampled large dataset can stand in for a small one. All 2,220 result files, the preregistration and its hash, and the code that regenerates every figure and number from those files are public.

发表机构

  • VIT Bhopal University(维特博帕尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑