表格基础模型蒸馏为高效预测器
Distillation of Tabular Foundation Models into Efficient Predictors
浏览论文内容
中文总结 AI 辅助
本研究针对表格基础模型推理成本高的问题,提出一种知识蒸馏方案,利用完整训练集作为教师上下文训练轻量级学生模型,在TabArena和TALENT上显著提升性能并实现3-21倍推理加速。
中文摘要 AI 辅助
表格基础模型(TFMs)通过上下文学习实现强大的预测性能,但反复依赖标注数据使得推理成本高昂。知识蒸馏可通过将预测能力转移至轻量级、特定数据集的学生模型来降低这一成本。然而,TFM预测对标注上下文和查询的双重依赖引出了两个设计问题:如何构建教师监督,以及扩展查询覆盖范围是否能改进蒸馏效果。我们针对两个TFM以及神经和基于树的学生模型研究了这些问题,并得出了一种有效的蒸馏方案。该方案使用完整标注训练集作为教师上下文,并仅基于教师对观测和合成查询的预测来训练学生模型。在TabArena上,所得学生模型在Elo评分上比其监督训练并调优集成的对应模型高出57-98分。将该方案原样应用于TALENT,在300个数据集中有236-258个数据集上改进了匹配的默认学生模型,并将中位主要误差降低了4.0-6.4%。蒸馏后的学生模型还实现了相对于教师模型3.0-21.6倍的中位推理加速,为预测性能与重复推理成本之间提供了实用的权衡。代码可在该https URL获取。
英文摘要
Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at https://github.com/nums-ai/TFM_Distillation .
发表机构
- Nums AI
机构由 AI 辅助整理,请以论文原文为准。