arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表格基础模型的测试时计算:机制、增益与局限

Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits

Kanghui Ning, Marin Biloš, James T. Wilson, Yilang Zhang, Kashif Rasul, Dongjin Song, Anderson Schneider, Yuriy Nevmyvaka

arXiv 2610.12005首次发表:更新:

发表机构

School of Computing, University of Connecticut; Morgan Stanley(康涅狄格大学计算机学院; 摩根士丹利)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从适应、聚合、上下文构建三方面探究测试时计算对表格基础模型的增益,提出DiagScale方法,发现适应与选择性聚合可获一致基准增益,需依计算预算选策略。

AI 中文摘要

哪些形式的测试时计算可提升强大的预训练表格基础模型(Tabular Foundation Models, TFMs)的预测性能?我们从适应、聚合和上下文构建三个维度对此展开系统研究。评估涵盖TabArena基准中的现代TFMs,并补充了OpenML上广泛且大规模表格的实验。对于适应,我们提出DiagScale,一种对角查询-键相似度更新方法,仅训练模型0.003%-0.03%的参数,在三个独立预训练的骨干模型上实现了与全微调相当的增益。对于聚合,池组成和选择策略均起作用:TabPFN-3已对同一数据的不同预处理变体的预测取平均,增加更多此类预测会产生边际收益递减;在包含96种配置的更广泛池中,贪心选择相比默认预测器降低了2.4%的误差,而均匀平均则会增加误差。对于上下文构建,注意力引导检索可提升TabPFN-3在部分大表格上的预测,并支持超出完整上下文记忆限制的源池;我们测试的上下文扩展方法未产生一致的改进。综上,我们的结果表明,适应和选择性聚合可产生一致的基准级增益;上下文构建的益处更多取决于任务和数据 regime;在相同骨干模型上结合适应与聚合可进一步提升增益,但相比默认推理需要显著更多计算;这些权衡促使根据可用计算预算选择策略。代码可在该https URL获取。

英文摘要

Which forms of test-time compute improve the predictions of strong pretrained tabular foundation models (TFMs)? We systematically study this along three axes: adaptation, aggregation, and context construction. Our evaluation spans modern TFMs across the TabArena benchmark, supplemented by experiments on wide and large-scale tables from OpenML. For adaptation, we introduce DiagScale, a diagonal query-key similarity update. It trains only 0.003-0.03% of model parameters and achieves gains comparable to full fine-tuning across three independently pretrained backbones. For aggregation, both pool composition and selection strategy matter. TabPFN-3 already averages predictions from different preprocessing variants of the same data, and adding more such predictions yields diminishing returns. With a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor, but uniform averaging increases error. For context construction, attention-guided retrieval improves TabPFN-3's predictions on some large tables and supports source pools beyond the full context memory limit. The context expansion methods we test yield no consistent improvement. Taken together, our results suggest that adaptation and selective aggregation yield consistent benchmark-level gains. The benefits of context construction depend more on the task and data regime. Adaptation and aggregation over the same backbone yield further gains when combined, but require substantially more computation than default inference. These trade-offs motivate choosing strategies according to the available computation budget. Code is available at https://github.com/kanghui-learning/test-time-compute-for-tabular-foundation-models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑