从良好起点到最优推断:带缺失和隐式正则化的广义潜因子模型
From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization
- Department of Statistics, Columbia University(哥伦比亚大学统计系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种结合链接感知双SVD初始化、单侧细化和梯度下降的非凸估计方法,实现广义潜因子模型在缺失数据下的最优推断与不确定性量化,无需显式正则化或去偏。
AI中文摘要:
广义潜因子模型为分析高维非高斯数据提供了灵活的框架,但在缺失情况下,基于原理的估计和不确定性量化仍远未发展成熟。我们发展了一套理论,将一种计算上易处理的非凸过程直接与具有指数族链接和部分观测条目的非线性潜因子模型的统计推断联系起来。我们的过程结合了链接感知的双奇异值分解初始化、实现行一致性的单侧细化以及朴素梯度下降。我们证明了细化后的初始化器进入非相干性和收缩区域,并且梯度下降通过隐式正则化保持在该区域内,在没有显式非相干性或平衡正则化的情况下快速收缩到统计估计误差。我们的核心结果是对梯度下降实际输出的均匀行向线性近似,该近似将主要得分波动与高阶估计和优化误差分离开来。这些展开式产生了渐近有效的个体推断和高斯乘子自助法同时推断潜因子,以及缺失条目均值的联合置信带,而无需额外的去偏步骤。由此产生的估计速率在对数因子范围内匹配受限类极小极大下界,而该理论适用于严重缺失、弱低秩信号和递减的局部曲率。模拟结果支持理论发现,对大型语言模型评估的应用展示了潜模型能力的不确定性感知估计和排名。
英文摘要:
Generalized latent factor models provide a flexible framework for analyzing high-dimensional non-Gaussian data, but principled estimation and uncertainty quantification under missingness remain substantially less developed. We develop a theory that connects a computationally tractable nonconvex procedure directly to statistical inference for nonlinear latent factor models with exponential-family links and partially observed entries. Our procedure combines a link-aware double-SVD initialization, a unilateral refinement that achieves rowwise consistency, and vanilla gradient descent. We show that the refined initializer enters a region of incoherence and contraction and that gradient descent remains in this region through implicit regularization, contracting rapidly down to the statistical estimation error without explicit incoherence or balancing regularization. Our central result is a uniform rowwise linear approximation for the actual output of gradient descent that isolates the leading score fluctuations from higher-order estimation and optimization errors. These expansions yield asymptotically valid individual and Gaussian multiplier-bootstrap simultaneous inference for latent factors, together with simultaneous confidence bands for missing-entry means, without requiring an additional debiasing step. The resulting estimation rate matches a restricted-class minimax lower bound up to logarithmic factors, while the theory accommodates severe missingness, weak low-rank signals, and diminishing local curvature. Simulations support the theoretical findings, and an application to large language model evaluation illustrates uncertainty-aware estimation and ranking of latent model capabilities.