arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机梯度下降中的可扩展统计推断

Scalable Statistical Inference in Stochastic Gradient Descent

Rahul Singh, Abhinek Shukla

arXiv 2608.30845首次发表:更新:

发表机构

Indian Institute Of Technology Delhi; Indian Institute of Technology Delhi-Abu Dhabi; Duke-NUS Medical School(印度理工学院德里分校; 印度理工学院德里-阿布扎比分校; 杜克-国大医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高维SGD置信区间构建的协方差估计瓶颈与退化问题,提出结合等批量均值、t分布边际统计量、wild bootstrap与t-copula拟蒙特卡罗的可扩展方法,经模拟验证可生成稳健高效的高维置信区间。

AI 中文摘要

为随机梯度下降(SGD)构建置信区间,理想情况下需要估计渐近协方差矩阵,这在高维场景中是严重的计算瓶颈。传统基于抵消的批量均值方法可绕过该估计,但需对样本批量协方差矩阵求逆,当参数维度超过批量数量时会引入严格的数学退化问题。为解决此问题,我们采用等批量大小的批量均值方法,提出一种同时性、边际友好型框架。所提边际统计量服从渐近学生t分布,消除了矩阵求逆步骤,完全规避了高维退化问题。为实现有效的同时覆盖,我们提出一种算法,利用仅基于方差-协方差估计器对角线函数的统计量生成的wild bootstrap样本,同时为纳入交叉依赖的贡献,引入利用t-copula近似的高效拟蒙特卡罗程序。此外,我们整合Lugsail方差估计器以积极修正有限样本偏差和覆盖不足问题。所提方法可生成可解释的同时性超矩形置信区间,具有统计稳健性、内存高效性,且对高维推断严格可扩展。理论结果通过维度、批量数量和误差结构等多方面的大量数值模拟分析得到支撑。

英文摘要

Constructing confidence regions for stochastic gradient descent (SGD) ideally requires estimating the asymptotic covariance matrix, a severe computational bottleneck in high dimensions. Traditional cancellation-based batch means methods bypass this estimation but require inverting a sample batch covariance matrix. This introduces strict mathematical degeneracy when the parameter dimension exceeds the number of batches. To address this problem, we utilize equal batch size batch means method and propose a simultaneous, marginal-friendly framework. The proposed marginal statistics has a asymptotic Student's $t$-distribution, and eliminates the matrix inversion step, entirely circumventing high-dimensional degeneracy. To achieve valid simultaneous coverage, we present an algorithm utilizing wild bootstrap samples drawn from a statistic as a function of only the diagonals of the variance-covariance estimator, and to further incorporate the contribution of cross-dependencies, we introduce an efficient Quasi-Monte Carlo procedure utilizing a $t$-copula approximation. Additionally, we integrate a Lugsail variance estimator to aggressively correct finite-sample bias and under-coverage. The proposed methodology delivers interpretable, simultaneous hyper-rectangular confidence regions that are statistically robust, memory-efficient, and strictly scalable for high-dimensional inference. The theoretical results are supported by extensive numerical simulation analysis through various aspects of dimension, number of batches and error structure.

Comments18 pages, 9 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑