发表机构
Georgia Institute of Technology; University of Southern California; University of Cambridge(佐治亚理工学院; 南加州大学; 剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对协变量缺失的逻辑回归,提出基于新颖单调算子的随机逼近方法,在协变量分布未知的假设精简设置下实现参数速率信号恢复,并优于完整案例估计器。
AI 中文摘要
在监督学习问题中,协变量缺失频繁出现,而利用此类数据进行估计的经典方法通常采用精心选择的缺失数据插补方案,或使用导致非凸$M$-估计问题的似然近似。这些方法及其变体适用于协变量分布已知的场景,更广泛地,它们在线性模型中取得了巨大成功。但即使在中等维度的逻辑回归等基本非线性问题中,当协变量分布未知时,这些方法也可能遭遇严重的失效模式。出于对可靠替代方案的需求,我们考虑逻辑回归中协变量缺失的参数估计问题。关键在于,我们在假设精简的环境中操作,其中协变量分布未知(但有界)。我们设计了一种基于新颖单调算子的$Z$-估计的随机逼近方法,并证明我们的算法在计算上高效,且在协变量完全随机缺失的假设下,能够以参数速率实现可证明的信号恢复。我们的理论以缺失概况精确刻画了估计量的$\ell_2^2$风险,适应了异质的观测概率。重要的是,它表明我们的方法总是优于忽略任何缺失数据观测的事实上的“完整案例”估计器。即使在同质缺失的设置中(每个协变量以概率$q$独立观测),我们的界也表现出对$q$的复杂且非标准的依赖,这可以带来比仅使用完整案例显著的改进。我们用新的信息论下界补充了上界,表明这种对$q$的复杂依赖在极小极大意义上是根本性的。
英文摘要
Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex $M$-estimation problems. These methods and their relatives are suitable for scenarios in which the covariate distribution is known, and more broadly, have enjoyed tremendous success in linear models. But even in basic nonlinear problems such as logistic regression in moderate dimensions, such methods can experience drastic failure modes when the covariate distribution is unknown. Motivated by the need for reliable alternatives, we consider the problem of parameter estimation in logistic regression with missing covariates. Crucially, we operate in the assumption-lean setting where the covariate distribution is unknown (but bounded). We design a stochastic approximation method that is based on $Z$-estimation with a novel monotone operator, and establish that our algorithm is computationally efficient and achieves provable signal recovery at parametric rates under the hypothesis that covariates are missing completely at random. Our theory sharply characterizes the $\ell_2^2$ risk of the estimator in terms of the missingness profile, accommodating heterogeneous observation probabilities. Importantly, it shows that our method always outperforms the de facto ``complete-case'' estimator that ignores observations with any missing data. Even in the setting with homogeneous missingness (in which each covariate is observed independently with probability $q$), our bounds exhibit intricate and nonstandard dependence on $q$ that can yield significant improvements over using only complete cases. We complement our upper bounds with new information-theoretic lower bounds that show that this intricate dependence on $q$ is fundamental in a minimax sense.