arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArborEnum:基于连续特征的决策树Rashomon集

ArborEnum: Decision Tree Rashomon Sets over Continuous Features

Zakk Heile, Hayden McTavish, Margo Seltzer, Cynthia Rudin

arXiv 2608.04310首次发表:更新:

AI 中文总结

ArborEnum算法可精确枚举基于连续特征的决策树Rashomon集,通过任意时间算法实现加速,相比现有方法有数量级提升,且近似版本能保持高召回率。

AI 中文摘要

Rashomon效应指的是在同一学习任务中,许多模型可达到近乎相当的性能,这对鲁棒性、特征重要性和可定制性有重大影响。这些应用场景促使人们计算Rashomon集:所有正则化损失接近最优的模型集合。决策树是少数能完全枚举Rashomon集的模型类别之一,但该计算始终依赖于对原始数据的二值化处理,要么限制每棵树允许进行的分裂操作,要么大幅增加这一已属困难的组合问题的复杂度。本文提出了首个能精确枚举决策树Rashomon集的算法,该算法可利用连续特征的有序结构。我们还开发了用于近似枚举的松弛方法,以及一种能逐步细化候选阈值集合的任意时间算法,该算法生成的近似结果会越来越详细,最终收敛到连续特征的Rashomon集。实验表明,粗糙的二值化会遗漏许多树、重要特征和预测多样性;与现有枚举方法相比,我们的算法实现了数量级的加速,而近似版本在保持近乎完美召回率的同时还能进一步提升速度。

英文摘要

The Rashomon effect describes the phenomenon that many models can achieve nearly equivalent performance on the same learning task, with significant ramifications for robustness, feature importance, and customizability. These use cases motivate the computation of Rashomon sets: the set of all models whose regularized loss is near-optimal. Decision trees are one of the few model classes for which Rashomon sets can be fully enumerated, but this computation has always been conditional on a binarization of the original data, either restricting which splits each tree is allowed to make or substantially increasing the complexity of an already difficult combinatorial problem. We introduce the first algorithm that exactly enumerates decision-tree Rashomon sets while exploiting the ordered structure of continuous features. We further develop a relaxation for approximate enumeration and an anytime algorithm that progressively refines the set of candidate thresholds, producing increasingly detailed approximations that converge to the continuous-feature Rashomon set. Experiments show that coarse binarization can miss many trees, important features, and predictive multiplicity; our algorithms achieve orders-of-magnitude speedups over existing enumeration methods, with approximations providing further speedups while maintaining near-perfect recall.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑