在线分位数回归的有限样本分布理论与高效大规模推断
Finite-Sample Distribution Theory and Efficient Large-Scale Inference for Online Quantile Regression
浏览论文内容
中文总结 AI 辅助
本文提出基于常数学习率SSGD的在线分位数回归推断方法,证明有限样本高斯近似,提出后缀平均化解决偏差,避免协方差估计,实现高效且覆盖率理想的推断。
中文摘要 AI 辅助
本文研究了使用常数学习率的随机次梯度下降(SSGD)进行大规模和流式数据的在线分位数回归。经典的离线分位数回归推断在计算和内存方面开销较大。现有的在线分位数回归推断工作仅提供渐近保证,并且通常需要次指数尾部条件来进行分布理论分析。为了弥合这些差距,我们引入了新技术,在有限矩假设下证明了SSGD的淬火中心极限定理(CLT)和有限样本高斯近似。我们进一步表明,具有常数学习率的Ruppert-Polyak平均化存在不可忽略的偏差,并且无法满足以总体目标为中心的CLT。因此,我们提出后缀平均化来解决此问题,并建立了其有限样本高斯近似。基于这些结果,我们提供了一种高效的分位数回归在线推断方法,该方法避免了协方差估计。数值实验表明,与其他推断方法相比,我们的方法实现了理想的经验覆盖率并具有竞争力的性能。我们还将我们的方法应用于美国工资数据,以证明其实际有效性。
英文摘要
This paper studies online quantile regression for large-scale and streaming data using Stochastic SubGradient Descent (SSGD) with constant learning rates. Classical offline inference for quantile regression is computationally and memory intensive. Existing works of online inference for quantile regression provide only asymptotic guarantees and typically require sub-exponential tail conditions for distribution theory. To bridge these gaps, we introduce new techniques to prove a quenched central limit theorem (CLT) and finite-sample Gaussian approximation for SSGD under a finite-moment assumption. We further show that Ruppert-Polyak averaging with a constant learning rate has a non-vanishing bias and fails to satisfy CLT centering at the population target. Hence we propose suffix averaging to address this issue and establish its finite-sample Gaussian approximation. Based on these results, we provide an efficient online inference method for quantile regression that avoids covariance estimation. Numerical experiments show that our method achieves desirable empirical coverage rates and competitive performance compared to other inference methods. We also apply our approach to U.S. wage data to demonstrate its practical effectiveness.
发表机构
- University of Chicago(芝加哥大学)
- Rice University(莱斯大学)
- University of Miami(迈阿密大学)
机构由 AI 辅助整理,请以论文原文为准。