AI 中文总结
该研究针对协变量偏移下的岭回归,对比池化与集成方法的性能,推导了固定和随机效应下的相关结论,推广了现有风险分析的适用场景。
AI 中文摘要
在许多场景中,数据集自然地划分为由亚群体、批次效应或多源聚合产生的簇。针对这种异质性,常见的应对方式是集成在每个簇上训练的学习器,而非对池化数据拟合单一模型。现有支持这类方法的研究通常考虑簇间协变量分布和条件结果模型均存在差异的场景;仅基于协变量分布的簇感知划分与集成的作用仍有待探索。我们针对线性结果模型下的岭正则化最小二乘回归处理该问题,考虑所有λ≥0的岭惩罚值,包括λ=0时的无岭预测器这一特殊情况。通过同时考虑固定效应和随机效应模型,我们证明在随机效应下,经最优调参的池化岭预测器始终优于经单独最优调参的预测器集成。对于固定效应,我们推导了池化和集成预测器的通用公式,以刻画回归系数及预测器分布偏移的作用。这些结果将基于岭和无岭回归预测器的Bagging及随机划分估计的现有风险分析,从独立同分布场景推广至包含协变量偏移和异质性感知划分结构的情况。
英文摘要
Datasets in many settings naturally partition into clusters arising from sub-populations, batch effects, or aggregation across multiple sources. A common response to such heterogeneity is to ensemble learners trained on each cluster rather than fit a single model to the pooled data. Prior work motivating such approaches has typically considered settings in which both the covariate distribution and the conditional outcome model differ across clusters; the role of cluster-aware partitioning and ensembling based solely on the covariate distribution remains to be explored. We address this case for ridge-regularized least-squares regression under a linear outcome model and consider all ridge penalty values $λ\geq 0$, including the special case of the ridgeless predictor at $λ= 0$. By considering both fixed-effects and random-effects models, we argue that under random effects, an optimally tuned pooled ridge predictor always outperforms ensembles of individually optimally tuned predictors. For fixed effects, we derive a general formula for the pooled and ensembled predictors to characterize the role of both regression coefficients as well as the predictor distribution shifts. Together, these results generalize prior risk analyses of bagging and random-partition estimation using ridge and ridgeless regression predictors from the i.i.d. setting to encompass covariate shift and heterogeneity-aware partition structure.
Comments9 pages, 2 figures