arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2610.04385math.OCcs.LG

用于鲁棒情境优化的最大Rashomon决策树集合

Largest Rashomon sets of decision trees for robust contextual optimization

  • Technical University of Denmark(丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Lorenzo Bonasera, David Pisinger

AI总结:

针对预测多重性下决策树拟合差异导致的决策不稳定问题,提出联合Rashomon与鲁棒性优化框架,通过约束生成算法高效求解,在报童和餐厅数据上显著提升鲁棒性并保持成本近似。

AI中文摘要:

许多决策树对同一数据的拟合效果几乎同样好,但它们可能将查询点路由到不同的叶节点,并诱导出不同的局部经验分布。我们研究在存在这种预测多重性的情况下,能够满足规定的成本、短缺或风险目标的决策。我们提出了针对最优决策树的联合Rashomon与鲁棒性优化框架。该框架联合选择一个操作决策和最大的Rashomon树集合,使得目标在该集合中每棵树在查询点诱导的局部经验分布下均成立。我们将该框架专门应用于回归设置,并证明树仅通过共享查询叶节点的训练观测影响决策,我们将这些观测称为其查询邻域。因此,鲁棒问题仅涉及有限个不同的约束,这些约束可以按估计损失递增的顺序进行检查。我们开发了一种约束生成算法,该算法将查询路径定价与动态规划相结合,以识别违反约束的邻域,而无需枚举所有树。在合成报童实例上,该算法通常只需要少量邻域,并且运行速度远快于完整邻域枚举。在餐厅需求数据上,相对于最优树的样本平均近似订单,鲁棒订单将平均可容忍的额外估计损失提高了17.6%,并将最差10%实际成本的经验条件风险值降低了8.3%,而平均成本差异在统计上不显著。可解释性分析进一步展示了保留的邻域如何解释决策及其鲁棒性极限。

英文摘要:

Many decision trees fit the same data almost equally well, yet they can route a query point to different leaves and induce different local empirical distributions. We study decisions that meet prescribed cost, shortage or risk targets despite this predictive multiplicity. We propose the joint Rashomon and robustness optimization framework for optimal decision trees. It jointly selects an operational decision and the largest Rashomon set of trees, so that the targets hold under the local empirical distribution that every tree in this set induces at the query point. We specialize the framework to the regression setting, and we show that a tree affects the decision only through the training observations sharing the query leaf, which we call its query neighborhood. As a result, the robust problem involves only finitely many distinct constraints, which can be examined in order of increasing estimation loss. We develop a constraint generation algorithm that combines query-path pricing with dynamic programming to identify violating neighborhoods without enumerating trees. On synthetic newsvendor instances, the algorithm typically needs few neighborhoods and runs substantially faster than full neighborhood enumeration. On restaurant demand data, the robust orders increase the mean tolerated excess estimation loss by 17.6% and reduce the empirical conditional value-at-risk of the worst 10% of realized costs by 8.3% relative to the sample average approximation orders of the optimal tree, while the mean cost difference is not statistically significant. An interpretability analysis further shows how the retained neighborhoods explain the decision and its robustness limit.

↑