随机森林的分布分裂准则:扩展、收缩与均值分裂的稳健性
Distributional Split Criteria for Random Forests: Extensions, Shrinkage, and the Robustness of Mean Splitting
浏览论文内容
中文总结 AI 辅助
研究随机森林的分布分裂准则,实现并研究了多个准则及森林权重收缩方法。通过多种比较刻画其适用情况,发现普通各向同性MMD较优,标量回归中均值分裂稳健,多变量响应时分布分裂作用明显,支持简单分配观点,相关内容在开源库中实现。
中文摘要 AI 辅助
分布随机森林用比较候选子节点中完整条件响应分布的准则取代了基于均值的CART分裂。我们在单一诚实森林实现中实现并系统研究了这类准则的一个家族:各向同性随机傅里叶特征最大均值差异(MMD)、各向异性对角带宽变体、自适应每次分裂频率选择变体和非核切片瓦瑟斯坦准则,以及森林权重的事后核均值收缩。通过对合成分位数机制、实际单变量基准、加利福尼亚住房子样本曲线以及多变量合成和实际响应进行成对种子比较,我们刻画了每种扩展的适用情况。有三个发现反复出现。首先,在分布准则中,普通各向同性MMD已接近同类最佳:各向异性、自适应频率和切片瓦瑟斯坦扩展以及事后收缩并没有系统地改进它。其次,在标量表格回归中,基于均值的CART分裂仍然是稳健的默认方法并在许多单元格中获胜。第三,多变量响应是分布分裂明显发挥作用的领域,在纯依赖copula上最为明显,即使边际CRPS没有区分,能量得分也能区分准则。证据支持一个简单的分配情况:只有当非位置结构既存在又可估计时,分布分裂才会有帮助;否则它会将分裂选择能力从均值稀释掉。所有准则、诚实森林和成对比较工具都在开源的\texttt{drforest}库中实现,其基于Rust的分裂搜索使广泛的准则扫描成本低廉。
英文摘要
Distributional random forests replace mean-based CART splitting with criteria that compare the full conditional response distribution in candidate children. We implement and systematically study a family of such criteria inside a single honest-forest implementation: isotropic random-Fourier-feature maximum mean discrepancy (MMD), an anisotropic diagonal-bandwidth variant, an adaptive per-split frequency-selection variant, and a non-kernel sliced-Wasserstein criterion, together with post-hoc kernel-mean shrinkage of the forest weights. Using paired-seed comparisons across synthetic quantile mechanisms, real univariate benchmarks, a California-housing subsample curve, and multivariate synthetic and real responses, we characterize where each extension pays. Three findings recur. First, among distributional criteria ordinary isotropic MMD is already close to best in class: the anisotropic, adaptive-frequency, and sliced-Wasserstein extensions, and post-hoc shrinkage, do not systematically improve on it. Second, on scalar tabular regression mean-based CART splitting remains the robust default and wins many cells. Third, multivariate responses are the regime where distributional splitting clearly earns its keep, most sharply on a pure-dependence copula where the energy score separates the criteria even though marginal CRPS does not. The evidence supports a simple allocation story: distributional splitting helps only when non-location structure is both present and estimable; otherwise it dilutes split-selection power away from the mean. All criteria, the honest forest, and the paired-comparison harness are implemented in the open-source \texttt{drforest} library, whose Rust-backed split search makes broad criterion sweeps inexpensive.