arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高维在线M估计的大规模线性假设检验

Large-scale linear hypothesis testing for high-dimensional online M-estimation

Daeyoung Ham, Myeonghun Yu, Tate Jacobson

arXiv 2610.07418首次发表:更新:

发表机构

University of Texas at San Antonio; Ewha Womans University; Oregon State University(德克萨斯大学圣安东尼奥分校; 梨花女子大学; 俄勒冈州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高维在线M估计,提出修正可再生损失结合部分折叠凹正则化,实现大规模线性假设检验,并建立误差界与渐近有效检验。

AI 中文摘要

随着流式数据的日益普及,可再生估计已成为在不保留历史原始观测值的情况下更新统计分析的重要工具。我们研究高维在线M估计中的估计与线性假设检验问题。现有的基于二次逼近的可再生程序可能产生一阶逼近误差,因为未惩罚损失的梯度在先前惩罚估计量处通常不消失,这对有效推断构成了特殊挑战。为解决此问题,我们提出一种修正的可再生损失,以保留历史一阶信息,并将其与部分折叠凹正则化及迭代局部线性逼近相结合。我们建立了有限样本的$\ell_1$-和$\ell_2$-误差界,以及约束和非约束迭代估计量的收缩界和强Oracle性质。基于这些结果,我们推导了Oracle Bahadur表示,并为一般线性假设建立了渐近有效的Wald检验和得分检验,允许检验Oracle维度和约束数量均发散。我们进一步刻画了它们在局部备择假设下的渐近功效。所提出的程序仅需递归更新的汇总统计量,同时恢复其合并数据Oracle对应物的一阶推断行为。模拟研究和实际数据应用证明了所提出的估计和检验程序的有限样本性能。

英文摘要

With the growing prevalence of streaming data, renewable estimation has become an important tool for updating statistical analyses without retaining historical raw observations. We study estimation and linear hypothesis testing in high-dimensional online M-estimation. Existing renewable procedures based on quadratic approximations can incur first-order approximation errors because the gradient of the unpenalized loss generally does not vanish at a preceding penalized estimator, posing a particular challenge for valid inference. To address this issue, we propose a corrected renewable loss that preserves historical first-order information and combine it with partial folded concave regularization and an iterative local linear approximation. We establish finite sample $\ell_1$- and $\ell_2$-error bounds, together with contraction bounds and strong oracle properties for both the constrained and unconstrained iterative estimators. Building on these results, we derive oracle Bahadur representations and establish asymptotically valid Wald and score tests for general linear hypotheses, allowing both the testing-oracle dimension and the number of restrictions to diverge. We further characterize their asymptotic power under local alternatives. The proposed procedure requires only recursively updated summaries while recovering the first-order inferential behavior of its pooled-data oracle counterpart. Simulation studies and real-data applications demonstrate the finite-sample performance of the proposed estimation and testing procedures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑