arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21866cs.LGstat.ML

表格数据上经典机器学习的缩放定律:一项基准研究

Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study

Kaihua Ding

首次发表
浏览论文内容

中文总结 AI 辅助

该研究对表格数据上经典机器学习进行课堂规模复制,学生对多个数据集和模型家族运行固定协议,得出幂律拟合情况、模型家族内近似共享指数及复制器实现方差三个发现,并发布相关数据。

中文摘要 AI 辅助

先前关于经典机器学习学习曲线的工作在小尺度上(通常是一条曲线、一个团队、少量单元格)将幂律拟合到表格数据上的树、线性和核模型。我们进行了一次分布式课堂规模的复制:127名学生每人对3个指定数据集运行固定协议,这些数据集来自18个表格分类和回归数据集以及6个模型家族(提升、随机森林、支持向量机、线性/逻辑、岭回归、套索回归),产生了11536次训练运行和1648条拟合的幂律曲线,形式为误差(N)=aN^(-b)+c。有三个发现:(1)幂律拟合:77.7%的单元格上R^2>0.8,在完整数据时树集成占主导(提升占50%的数据集,随机森林占33%;线性模型在分类上表现不佳)。(2)模型家族内近似共享指数:6个家族中的5个,单个家族级指数预测每个家族的跨数据集曲线几乎与每个数据集的指数一样好(R^2差距<0.011),尽管AIC支持无约束拟合且曲线坍缩是部分的(32 - 58%的点在+/-0.5 dex内)。我们将此视为近似预测可压缩性,而非与数据集无关的普遍性;套索回归完全失败(负控制)且岭回归在留一数据集时很脆弱。(3)复制器实现方差:在固定random_state=42时,相同协议的独立重新实现对拟合指数的平均CV(b)=0.144仍有差异——不是种子方差,而是协议无约束部分(预处理、编码、缺失值处理)引起的差异。我们发布了聚合曲线、每个单元格的拟合以及达到目标误差0.15时N*的实用数据需求表。

英文摘要

Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We present a distributed classroom-scale replication: 127 graduate students each ran a fixed protocol on 3 assigned datasets, drawn from 18 tabular classification and regression datasets and 6 model families (Boosting, Random Forest, SVM, Linear/Logistic, Ridge, Lasso), yielding 11,536 training runs and 1,648 fitted power-law curves of the form error(N) = a N^(-b) + c. Three findings. (1) Power laws fit: R^2 > 0.8 on 77.7% of cells, with tree ensembles dominating at full data (Boosting 50% of datasets, RandomForest 33%; linear models underperform on classification). (2) Approximate shared exponents within a model family: for 5 of 6 families, a single family-level exponent predicts each family's cross-dataset curves nearly as well as per-dataset exponents (R^2 gap < 0.011), though AIC favors the unconstrained fit and curve collapse is partial (32-58% of points within +/-0.5 dex). We frame this as approximate predictive compressibility, not dataset-independent universality; Lasso fails outright (negative control) and Ridge is fragile under leave-one-dataset-out. (3) Replicator-implementation variance: with random_state=42 fixed, independent re-implementations of the same protocol still differ by mean CV(b) = 0.144 on the fitted exponent -- not seed variance, but the spread induced by unconstrained parts of the protocol (preprocessing, encoding, missing-value handling). We release the aggregated curves, per-cell fits, and a practical data-requirement table for N* to reach target error 0.15.

发表机构

  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑