发表机构
Chalmers University of Technology; University of Gothenburg(查尔姆斯理工大学; 哥德堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出表格Transformer模型MotherTree,通过合成数据元学习生成独立决策树,在基准测试中优于从头学习的梯度方法,可作为强初始化器提升决策树训练效果。
AI 中文摘要
传统决策树算法可生成有效、透明的模型,这类模型可被审计、交流,且能独立于训练数据部署,但需从头学习每个新任务。相比之下,表格基础模型表明,从合成先验分布进行元学习可对未见任务实现强上下文预测,尤其在小样本场景中;不过该方法无法生成可单独检查的独立模型。我们提出MotherTree,这是一种表格Transformer,可对决策树归纳进行元学习:给定新任务的训练集,它在单次前向传播中输出硬轴对齐决策树,其形式与经典训练的决策树等价。MotherTree使用随机梯度下降在合成先验上预训练,无需参考树作为监督。在样本量可控的既定基准上,该方法与常见算法生成的规模匹配的决策树具有竞争力,这些算法包括递归分区、基于梯度的树学习、全局最优树以及从表格基础模型进行蒸馏。值得注意的是,MotherTree始终优于从头开始的基于梯度的学习,且可作为强初始化器:在所有基准和样本量下,对生成树进行特定任务调优的表现均优于对应的从头学习器。这些结果表明,元学习可为学习独立的小型决策树分类器提供有效的归纳偏置。
英文摘要
Conventional decision tree algorithms produce effective, transparent models that can be audited, communicated, and deployed independently of the training data, but require learning every new task from scratch. In contrast, tabular foundation models demonstrate that meta-learning from a synthetic prior distribution enables strong in-context prediction for previously unseen tasks, especially in small-sample regimes. However, this approach does not produce a standalone model that can be inspected in isolation. We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass. MotherTree is pre-trained on a synthetic prior using stochastic gradient descent without requiring reference trees for supervision. On established benchmarks with controlled sample size, the approach is competitive with size-matched trees from common algorithms: recursive partitioning, gradient-based tree learning, globally optimal trees, and distillation from tabular foundation models. Notably, MotherTree consistently improves over from-scratch gradient-based learning and acts as a strong initializer: task-specific tuning of the generated tree outperforms the corresponding from-scratch learner on all benchmarks and sample sizes. These results show that meta-learning can provide effective inductive biases for learning stand-alone, small decision tree classifiers.