在LLM时代训练经典模型是否仍然值得?关于表格数据的交叉基准研究
Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data
查看机构详情
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过交叉点N*量化比较冻结LLM与经典模型在表格数据上的性能,发现训练经典模型在86%情况下用现有数据即可胜出,建议收集数百标签训练梯度提升模型。
中文摘要 AI 辅助
大型语言模型可以从纯英文描述中标记表格行而无需训练——这一能力现已出现在主流电子表格工具中,如Microsoft Copilot in Excel和Anthropic的Claude for Excel——这为许多标签昂贵的商业预测问题提出了一个实际问题:你应该提示一个冻结的LLM,还是收集数据并训练一个模型——如果是这样,需要多少数据?我们通过带标签数据的交叉点N*来量化答案,即训练好的经典模型的学习曲线超过冻结LLM的无训练(因此平坦)误差时的训练集大小。汇总了在18个表格数据集上、八种提示配置下对小型GPT模型的126项独立学生评估,并结合六个经典模型族的权威幂律学习曲线,我们发现训练胜出很快:即使给定其最佳提示配置的预言机选择,训练好的经典模型在86%的情况下使用不超过手头已有的带标签数据就能击败小型冻结LLM,在40%的情况下通过我们评估的最小带标签子集获胜,观察到的交叉点中位数约为训练集的6%。上下文中的少样本示例行为不像训练——误差与样本数量的关系不遵循幂律——并且由独立实施者重新运行的相同协议以0.148的变异系数变化。一项受控探针表明LLM依赖于可识别的特征名称语义,这可能使我们的交叉点成为保守估计(我们不声称记忆)。对于典型的商业表格,证据是明确的:收集几百个标签并训练一个梯度提升模型。
英文摘要
Large language models can label a tabular row from a plain-English description with no training - a capability now shipping in mainstream spreadsheet tools such as Microsoft Copilot in Excel and Anthropic's Claude for Excel - raising a practical question for the many business prediction problems where labels are expensive: should you prompt a frozen LLM, or collect data and train a model - and if so, how much data? We quantify the answer with the labeled-data crossover N*, the training-set size at which a trained classical model's learning curve overtakes a frozen LLM's training-free (and therefore flat) error. Aggregating 126 independent student evaluations of small GPT models under eight prompting configurations across 18 tabular datasets, paired with authoritative power-law learning curves for six classical model families, we find that training wins fast: even given an oracle choice of its best prompt configuration, a trained classical model beats the small frozen LLM using no more labeled data than is already on hand in 86% of cases, and wins by the smallest labeled subset we evaluate in 40%, with the observed crossover at a median of ~6% of the training set. In-context few-shot examples do not behave like training - error versus shot count does not follow a power law - and the same protocol re-run by independent implementers varies with a coefficient of variation of 0.148. A controlled probe indicates the LLM depends on recognizable feature-name semantics, which plausibly makes our crossover a conservative estimate (we do not claim memorization). For a typical business table, the evidence is clear: collect a few hundred labels and train a gradient-boosted model.