发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Literati是首个通过AO*算法实现最优形状广义树归纳的方法,联合优化树结构和形状函数复杂度,在24个数据集上超越现有树方法。
AI 中文摘要
决策树因其可解释性和在表格数据上的强大性能而备受青睐,但流行的贪心自顶向下归纳算法可能产生次优且不必要的复杂结构。最优决策树方法通过全局优化解决这一问题,但仍局限于轴对齐的阈值分割,这限制了每个节点的表达能力,并常常迫使深层复杂的树来捕捉非线性特征效应。形状广义树(SGTs)将阈值分割推广为可学习的单变量形状函数,提高了表达能力并使得树更加紧凑。然而,现有的SGT归纳算法是贪心的,不提供最优性保证。在这项工作中,我们引入了Literati,这是第一个用于最优SGT归纳的算法。我们提出了一种新颖的AND/OR图问题表述,该表述联合优化树结构和形状函数复杂度。为了解决这个AND/OR图,我们开发了一种基于AO*的算法,并带有两项增强功能,在保持最优性的同时提高任意时刻性能:用于OR节点选择的次级启发式策略和用于AND节点探索的轮询策略。在24个真实世界数据集上,Literati实现了比最先进的树方法更高的训练和测试准确率。
英文摘要
Decision trees are prized for their interpretability and strong performance on tabular data, but popular greedy top-down induction algorithms can yield suboptimal and unnecessarily complex structures. Optimal decision tree methods address this through global optimization, yet remain restricted to axis-aligned threshold splits, which limit the expressivity of each node and often force deep, complex trees to capture non-linear feature effects. Shape Generalized Trees (SGTs) generalize threshold splits to learnable univariate shape functions, improving expressivity and enabling more compact trees. However, existing SGT induction algorithms are greedy and offer no optimality guarantees. In this work, we introduce Literati, the first algorithm for optimal SGT induction. We propose a novel AND/OR graph formulation of the problem that jointly optimizes tree structure and shape function complexity. To solve this AND/OR graph, we develop an AO*-based algorithm with two enhancements that improve anytime performance while preserving optimality: a secondary heuristic for OR-node selection and a round-robin policy for AND-node exploration. Across 24 real-world datasets, Literati achieves higher training and test accuracy than state-of-the-art tree approaches.