用于XGBoost的卷积平滑分位数回归
Convolution Smoothed Quantile Regression for XGBoost
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究开发了基于分位数的梯度提升框架QXGB,引入卷积平滑损失,通过模拟数据和PM$_{2.5}$预测应用验证其在分位数估计、尾部表征等方面的优势。
AI中文摘要:
许多科学领域中大量复杂数据集的日益可得,使得机器学习(ML)被广泛应用于预测。然而,大多数ML算法聚焦于点估计,仅能提供有限的预测不确定性或响应条件分布信息,限制了其对稀有或极端结果的表征能力。我们开发了QXGB,一种基于分位数的梯度提升框架,并在其中引入卷积平滑损失,用于估计条件分位数,以构建与极端结果相关的密集累积分布函数(CDF)、超越概率及尾部行为。该方法在保留极端梯度提升计算效率的同时,恢复了XGBoost用于树分裂所依赖的海森信息,进而提供可解释的极值及超越概率预测度量。我们推导了将卷积平滑分位数损失与不同核规格整合至XGBoost所需的梯度和海森矩阵,并通过模拟数据,将该方法与其他平滑分位数回归损失、XGBoost Python包中的原生分位数目标,以及独立与多输出树估计进行基准测试。我们通过预测北加利福尼亚州细颗粒物(PM$_{2.5}$)的应用,包括野火烟雾导致浓度升高的时期,说明了该方法的实际相关性。结果表明,卷积平滑QXGB(尤其与多输出树配对时)能提供准确预测,且分位数交叉接近零,CDF和超越概率估计校准良好,对极端值的尾部表征有效,还评估了区间估计作为数据分布的度量。
英文摘要:
The increasing availability of large and complex datasets across many scientific disciplines has led to widespread adoption of machine learning (ML) for prediction. However, most ML algorithms focus on point estimation and provide limited information about predictive uncertainty or the conditional distribution of the response, restricting their ability to characterize rare or extreme outcomes. We develop QXGB, a quantile-based gradient boosting framework, and introduce a convolution smoothed loss within it that estimates conditional quantiles for constructing dense cumulative distribution functions (CDFs), exceedance probabilities, and tail behaviour relevant to extreme outcomes. This approach preserves the computational efficiency of extreme gradient boosting while restoring the Hessian information XGBoost relies on for tree splitting, in turn providing interpretable measures of extreme value and exceedance probability predictions. We derive the gradients and Hessians needed to integrate convolution smoothed quantile loss with different kernel specifications into XGBoost, and with simulated data, benchmark this approach against alternative smoothed quantile regression losses, the native quantile objective in the XGBoost Python package, and independent versus multi-output tree estimation. The practical relevance is illustrated in an application predicting fine particulate matter (PM$_{2.5}$) in northern California, including periods where levels were elevated due to wildfire smoke. Our results show that convolution smoothed QXGB, particularly when paired with multi-output trees, delivers accurate predictions with near-zero quantile crossing, well-calibrated CDF and exceedance probability estimates, and useful tail characterization for extreme values. Interval estimation is also evaluated as a measure of data spread.