arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05826cs.LG

通过自适应特征与样本缩减扩展最优分类树的规模

Scaling Optimal Classification Trees via Adaptive Feature and Sample Reduction

  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiancheng Tu, Wenqi Fan

AI总结:

针对最优分类树动态规划计算昂贵的问题,提出基于STreeD的联合特征与样本缩减框架,通过加权合并与自适应候选集优化,实现最高121倍加速并保持预测性能。

AI中文摘要:

动态规划用于最优分类树时,随着特征数和训练样本数的增加,计算成本变得非常高。我们基于STreeD开发了一种联合特征空间与样本空间的缩减框架。加权STreeD将投影到固定候选集后产生的重复记录合并为加权代表,从而在不改变固定候选优化问题的情况下减少依赖样本的计算量。自适应STreeD反复细化有界候选集,保留当前最优树使用的特征,重建加权表示,并求解由此产生的缩减问题。每个经过认证的加权STreeD解对其当前候选集都是最优的,而外层特征搜索在整个特征空间上仍是启发式的。在五个数据集上的实验表明,加权STreeD相比标准STreeD实现了最高121.41倍的加速。自适应STreeD在深度2至4的匹配比较中减少了运行时间,并在更深的深度下(全特征方法受时间或内存限制)继续返回可行树。在相同的计算预算下,其预测性能与所评估的最优分类树基线相当,并在某些比较中更高。这些结果表明,联合特征与样本空间缩减可以将基于动态规划的最优树学习扩展到更具挑战性的实例。

英文摘要:

Dynamic programming for optimal classification trees becomes computationally expensive as the numbers of features and training samples increase. We develop a joint feature- and sample-space reduction framework based on STreeD. Weighted STreeD merges duplicate records created after projection onto a fixed candidate set into weighted representatives. This reduces sample-dependent computation without changing the fixed-candidate optimization problem. Adaptive STreeD repeatedly refines a bounded candidate set, retains features used by the incumbent tree, rebuilds the weighted representation, and solves the resulting reduced problems. Each certified Weighted STreeD solution is optimal for its current candidate set, while the outer feature search remains heuristic over the full feature space. Experiments on five data sets show that Weighted STreeD achieves speedups of up to 121.41 times over standard STreeD. Adaptive STreeD reduces runtime in matched comparisons at depths 2 to 4 and continues to return feasible trees at greater depths where full-feature methods are limited by time or memory. Under the same computational budget, its predictive performance remains comparable to the evaluated optimal classification tree baselines and is higher in some comparisons. These results show how joint feature- and sample-space reduction can scale dynamic-programming-based optimal-tree learning to more demanding instances.

↑