arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GEAR:面向表格基础模型两阶段蒸馏的生成式扩展与真实锚定

GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models

Qi Qin, Jiajie Zhu, Dali Chen, Yuzhao Zhang, Jia-Xing Han, Peng Zhang, Ying Yan, Yifan Sun, Yu Su

arXiv 2608.18849首次发表:更新:

AI 中文总结

GEAR是一种两阶段蒸馏框架,可将表格基础模型蒸馏为轻量级预测器,在TALENT和TabArena上的实验显示其能显著降低推理开销并提升AUC,性能优于多种基准模型。

AI 中文摘要

表格基础模型(Tabular Foundation Models, TFMs)通过上下文学习实现了优异性能,但依赖上下文的推理会带来显著的延迟和内存开销,阻碍了其大规模部署。我们提出了GEAR(Generative Expansion and Real Anchoring,生成式扩展与真实锚定),这是一种模块化的两阶段框架,可将TFMs蒸馏为可在商用CPU上部署的轻量级MLP或基于树的预测器。第一阶段仅使用合成协变量作为教师查询位置,并在TFMs的软标签上训练学生模型,从而扩展了观测行之外的覆盖范围。第二阶段利用真实标签和折外教师预测重新将学生模型锚定到目标分布,避免了自标记泄漏。我们进一步推导了风险证书,用于表征生成查询量与生成器保真度之间的权衡。在TALENT和TabArena上的实验证明了GEAR的广泛适用性。两阶段MLP在二分类任务上比监督式MLP高出1.81-2.00个AUC点,在多分类任务上高出1.19-1.35个点;相比仅使用真实数据的蒸馏方法,额外提升分别为1.76-2.19和2.09-2.40个点。在二分类任务上,这些提升也可迁移至LightGBM和XGBoost,且三类学生模型在平均AUC上均优于最强的非TFM基准模型CatBoost。 ablation实验显示,其收益超出更长时间训练或替代预热启动的效果;分阶段优化比混合优化具有更高的稳定性;随着查询量增加,收益会随生成器变化而呈现边际递减。最后,GEAR将中位数推理时间降低了57-2866倍,峰值预测内存降低了1.9-3.3倍,同时保留了比匹配的监督式基准模型更高的AUC。

英文摘要

Tabular foundation models (TFMs) achieve strong performance through in-context learning, but context-dependent inference imposes substantial latency and memory costs, hindering large-scale deployment. We propose GEAR (\emph{Generative Expansion and Real Anchoring}), a modular two-stage framework that distills TFMs into lightweight MLP or tree-based predictors that can be deployed on commodity CPUs. Stage 1 uses synthetic covariates solely as teacher-query locations and trains the student on soft TFM targets, expanding coverage beyond observed rows. Stage 2 re-anchors the student to the target distribution using real labels and out-of-fold teacher predictions, whitch avoids self-labeling leakage. We further derive a risk certificate characterizing the trade-off between generated-query volume and generator fidelity. Experiments on TALENT and TabArena demonstrate the broad applicability of GEAR. Two-stage MLPs outperform supervised MLPs by 1.81--2.00 AUC points on binary tasks and 1.19--1.35 points on multiclass tasks, with additional gains over real-data-only distillation of 1.76--2.19 and 2.09--2.40 points, respectively. On binary tasks, the gains also transfer to LightGBM and XGBoost, and all three student families outperform CatBoost, the strongest non-TFM baseline, in mean AUC. Ablations show gains beyond longer training or alternative warm starts, greater stability from staged than mixed optimization, and generator-dependent diminishing returns as query volume increases. Finally, GEAR reduces median inference time by 57--2866 times and peak prediction memory by 1.9--3.3 times, while retaining higher AUC than matched supervised baselines.

Comments9 pages,5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑