通过团搜索识别非线性和依赖的潜在因子结构
Identification of Nonlinear and Dependent Latent Factor Structure through Clique Search
- UCLA Department of Statistics & Data Science(加州大学洛杉矶分校统计与数据科学系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对潜在因子模型的结构学习难题,提出基于成对依赖性度量的团搜索方法,可同时识别因子数量与非线性映射,并借助依赖性阈值化算法和神经网络实现准确估计,具备鲁棒性。
AI中文摘要:
学习潜在因子模型的结构涉及两个核心挑战:(1)估计潜在因子的数量,(2)学习从潜在变量到观测变量的映射的支持集。这在非参数体制和非线性设置中尤其具有挑战性。我们提出了一种基于观测变量上的成对依赖性度量、利用图论表示的潜在结构学习方法。我们表明,在温和的结构假设下,潜在因子的数量和非线性映射结构都可以从观测变量的分布中识别出来。与先前仅限于线性相关的工作不同,我们在非线性因子模型下为一般类别的依赖性度量建立了可识别性和一致性。这催生了一种依赖性阈值化(DT)算法,该算法仅从观测数据中联合估计潜在因子的数量和非线性映射结构。我们将其与受非线性映射结构约束的神经网络架构配对,以恢复非线性函数。通过模拟研究,我们表明DT算法在实践中是准确的,即使使用诸如神经网络之类的灵活方法,并且在高维设置中对其假设的违反表现出鲁棒性。
英文摘要:
Learning the structure of latent factor models involves two central challenges: (1) estimating the number of latent factors and (2) learning the support of the mapping from latent variables to observed variables. This is especially challenging for nonparametric regimes and nonlinear settings. We propose a method for latent structure learning, based on pairwise dependence measures on the observed variables using a graph-theoretic representation. We show that both the number of latent factors and nonlinear mapping structure can be identified from the distribution of observed variables under mild structural assumptions. Unlike prior work restricted to linear correlations, we establish identifiability and consistency for a general class of dependence measures under nonlinear factor models. This motivates a Dependence Thresholding (DT) algorithm, which jointly estimates the number of latent factors and nonlinear mapping structure from observational data alone. We pair this with a neural network architecture constrained by the nonlinear mapping structure, to recover the nonlinear function. Through simulation studies, we show that the DT algorithm is accurate in practice, even when using flexible methods such as neural networks, and exhibits robustness against violations of its assumptions in high-dimensional settings.