发表机构
Fundamental Technologies; Voylab(基础技术; Voylab)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究通用大型语言模型(LLM)在表格预测任务中表现不佳的原因,通过控制实验排除了数据噪声、CSV格式等因素,发现维度是关键,LLM准确率随维度增长下降,而经典基线方法不受影响。
AI 中文摘要
大型语言模型(LLMs)已成为众多任务的默认工具,但在最常见的机器学习工作负载之一——表格数据的预测分析中,却几乎没有取得成功。这一差距是快速发展的表格基础模型领域的创立前提,但通用LLMs为何失败的问题仍未得到解答。我们在纯推理场景下研究一款前沿LLM:仅通过一次生成处理包含全部训练和测试数据的提示,不借助任何工具、智能体框架,也不进行微调,系统评估了导致失败的五种假设:(a) 无法处理噪声或非线性可分数据;(b) 线性化的CSV格式会掩盖列结构;(c) 数值的分词处理;(d) 每次查询分类的测试点数量;(e) 输入的维度。受控实验否定了假设(a)-(d),而维度则是决定性因素:对31个基准数据集进行大范围随机线性投影后,在9种方法中,只有该LLM的准确率随维度增长而下降,而所有经典基线方法的准确率保持稳定或提升。与252种配置后的经典模型进行行为对比发现,在二维场景下,该LLM的预测类似基于局部距离的方法(网格一致性最高达91.6%),但在更高维度下,即使是配备了调优后与维度相关噪声的经典模型,也无法复现其预测结果。我们并未声称已确定内部机制;更严谨地说,我们的结果表明,该LLM的能力随维度增长而消解,这是任何受噪声干扰的经典学习器都无法模拟的——这解释了为何LLMs在其他领域表现出色,却在表格任务上持续输给已有五十年历史的基线方法,而预测的内部机制仍是一个悬而未决的问题。
英文摘要
Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remained open. We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning - and systematically evaluate five hypotheses for the failure: (a) an inability to handle noisy or non-linearly-separable data; (b) the linearised CSV format obscuring column structure; (c) the tokenisation of numeric values; (d) the number of test points classified per query; and (e) the dimensionality of the input. Controlled experiments falsify (a)-(d). Dimensionality, in contrast, is decisive: sweeping random linear projections of thirty-one benchmark datasets, the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. A behavioural comparison against 252 configured classical models finds that in two dimensions the LLM predicts like a local, distance-based method (up to 91.6% grid agreement), but in higher dimensions no classical model - even when augmented with tuned, dimension-dependent noise - reproduces its predictions. We do not claim to have identified the internal mechanism; our results show, more modestly, that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.