arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

处理概率回归树中的缺失数据

Handling Missing Data in Probabilistic Regression Trees

Taiane Schaedler Prass, Alisson Silva Neimaier, Guilherme Pumi

arXiv 2608.06195首次发表:更新:

发表机构

Universidade Federal do Rio Grande do Sul(南里奥格兰德联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文扩展PRTree框架,提出三种处理缺失数据的策略,经多数据集验证,所提方法在缺失比例高时优于CART,且保留树模型的可解释性与灵活性。

AI 中文摘要

概率回归树(PRTrees)是经典回归树的平滑且一致的替代方案,通过概率分裂分配生成连续预测结果。本文将PRTree框架扩展为可在树构建过程中直接处理预测变量缺失值,无需预先插补。提出三种策略,每种策略利用可用信息的方式不同:均匀概率法、部分观测法和降维平滑法。这些修改旨在在协变量值存在任意缺失模式时,保留原始方法的基本概率属性,包括概率守恒和边际兼容性。在几个具有不同缺失程度的真实世界数据集上对所提方法进行评估,并与经典回归树进行比较。结果表明,概率树构建的有效性在很大程度上取决于对缺失观测值的处理。在考虑的数据集中,填充策略成为主要建模组件,通常对预测性能的影响大于平滑分布或代理选择准则。在相当比例的观测值包含预测变量缺失值的数据集中,所提方法经常优于CART,同时保持基于树模型的可解释性和灵活性。

英文摘要

Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the need for prior imputation. Three strategies are proposed, each exploiting the available information differently: a uniform-probability approach, a partial-observation approach, and a dimension-reduced smoothing approach. These modifications are defined to preserve the fundamental probabilistic properties of the original methodology, including probability conservation and marginal compatibility, under arbitrary patterns of missing covariate values. The proposed methods are evaluated on several real-world datasets exhibiting different levels of missingness and are compared with classical regression trees. The results show that the effectiveness of probabilistic tree construction depends strongly on the treatment of missing observations. Across the considered datasets, the fill strategy emerged as the dominant modeling component, often exerting a larger influence on predictive performance than either the smoothing distribution or the proxy-selection criterion. In datasets where a substantial proportion of observations contained missing predictor values, the proposed methods frequently outperformed CART, while maintaining the interpretability and flexibility of tree-based models.

CommentsTheoretical background for the companion paper, arXiv:2510.03634

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑