AI 中文总结
本文针对高斯设计下逻辑回归的参数估计问题,构造了首个极小极大最优估计器,改进了MLE的有限样本误差率,通过数值实验验证了所提估计器性能优于MLE。
AI 中文摘要
我们研究高斯设计下逻辑回归的有限样本参数估计问题,目标是从独立同分布样本 $\{(\mathbf{x}_i,y_i)\}_{i=1}^n$ 中估计 $\mathbf{\theta}^*\in \mathbb{R}^d$,其中 $R=\\|\mathbf{\theta}^*\\|_2\ge 1$,$\mathbf{x}_i \sim N(0,\mathbf{I}_d)$,$y_i\mid \mathbf{x}_i \sim \mathrm{Bernoulli}((1+\exp(-\mathbf{x}_i^\top \mathbf{\theta}^*))^{-1})$。本文中,我们提供了首个极小极大最优估计器,并改进了极大似然估计(MLE)的已知最优有限样本误差率。这两项成果均源于对参数范数 $R$ 的极小极大最优估计器的研究。首先,我们建立了范数估计的极小极大下界 $\Omega(\sqrt{R^3/n})$。随后,我们将 Chardon、Lerasle 和 Mourtada(2024)给出的 MLE 的最优已知范数估计误差率 $O(\sqrt{R^3d/n})$ 改进为 $\tilde{O}(\sqrt{R^3/n}+R^2d/n)$。正如 Zhao、Sur 和 Candes(2022)的高维渐近理论及数值例子所证实,额外项 $R^2d/n$ 似乎是 MLE 的固有偏差。然而,我们表明该额外项并非信息论上必需的。为此,我们构造了一种高效的去偏范数估计器,其误差率达到 $O(\sqrt{R^3/n})$,因此是极小极大最优的。将此与 MLE 给出的最优方向估计器相结合,我们建立了估计 $\mathbf{\theta}^*$ 的极小极大最优率 $\Theta(\sqrt{Rd/n}+\sqrt{R^3/n})$,以及 MLE 的改进有限样本误差率 $\tilde{O}(\sqrt{Rd/n}+\sqrt{R^3/n}+R^2d/n)$。数值实验表明,所提出的极小极大最优估计器的性能优于 MLE。
英文摘要
We study finite-sample parameter estimation in logistic regression with Gaussian design, where the goal is to estimate $\mathbfθ^*\in \mathbb{R}^d$ with $R=\|\mathbfθ^*\|_2\ge 1$ from i.i.d. samples $\{(\mathbf{x}_i,y_i)\}_{i=1}^n,$ $\mathbf{x}_i \sim N(0,\mathbf{I}_d)$, $y_i\mid \mathbf{x}_i \sim \mathrm{Bernoulli}((1+\exp(-\mathbf{x}_i^\top \mathbfθ^*))^{-1})$. In this paper, we provide the first minimax optimal estimator, and improve on the best known finite-sample error rate for the maximum likelihood estimator (MLE). These two accomplishments are due to a minimax optimal estimator for the parameter norm $R$. First, we establish the minimax lower bound $Ω(\sqrt{R^3/n})$ for norm estimation. We then improve the best known norm estimation error rate of the MLE, i.e., $O(\sqrt{R^3d/n})$ from Chardon, Lerasle and Mourtada (2024), to $\tilde{O}(\sqrt{R^3/n}+R^2d/n)$. The additional term, $R^2d/n$, appears to be the intrinsic bias of the MLE, as evidenced by the high-dimensional asymptotic theory of Zhao, Sur and Candes (2022) and numerical examples. We show that, however, this additional term is not information-theoretically necessary. To this end, we construct an efficient debiased norm estimator that achieves the error rate $O(\sqrt{R^3/n})$ and is therefore minimax optimal. Combining this with the optimal direction estimator given by the MLE, we establish the minimax optimal rate $Θ(\sqrt{Rd/n}+\sqrt{R^3/n})$ for estimating $\mathbfθ^*$, as well as the improved finite-sample error rate $\tilde{O}(\sqrt{Rd/n}+\sqrt{R^3/n}+R^2d/n)$ for the MLE. Numerical experiments demonstrate that the proposed minimax optimal estimators outperform the MLE.