高维网络与可能存在误设定模型的均方误差
High-dimensional networks and mean squared error for possibly misspecified models
- University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文针对高维网络场景,研究了可能存在误设定模型的均方误差,证明最小描述长度方法可实现低假阳性率的邻域选择,颠覆了传统模型选择观点。
中文摘要 AI 辅助
为避免在网络分析中遗漏重要变量及其关联,越来越多的变量被纳入网络分析。本文表明,在参数数量远多于观测值(高维)的场景下,能够对每个节点的邻域(网络中的连接)给出保守(即低假阳性率)的估计。邻域通常通过线性模型估计,这会产生两种有趣的情况:(i)若真实模型为线性,则邻域选择效果较好;(ii)若真实模型为非线性,则邻域选择需要对高维情况施加惩罚。本文阐述了岭参数对均方误差的影响,以及这如何导致低测试方差,进而形成边数较多的邻域。我们将这些见解与机器学习领域的结果相联系,其中所谓的双重下降(当参数数量超过观测值时,均方误差会再次下降)颠覆了传统的模型选择观点。本质上,为了在参数数量庞大的模型中实现合适的邻域选择,惩罚项需要包含模型空间的体积。大多数邻域选择方法(如Lasso、AIC、BIC)会产生虚假边(高假阳性率),但我们证明,在高维场景下,最小描述长度方法无论模型是否被误设为线性,都能实现正确的邻域选择或更低的假阳性率。
英文摘要
To avoid missing important variables and their connections in networks, more and more variables are included in network analysis. Here we show that in a setting with many more parameters than observations (high-dimensional) it is possible to get a conservative (i.e., low false positive rate) estimate of the neighbourhood for each node (which connections are in the network). A neighbourhood is often estimated with a linear model, and this leads to two interesting cases: (i) If the true model is linear, then neighbourhood selection work reasonably well, and (ii) if the true model is nonlinear, then neighbourhood selection requires a penalty for the high dimensions. Here we show the impact of the ridge parameter on the mean squared error, and how this leads to low test variance and hence to neighbourhoods with large numbers of edges. We connect these insights with results from machine learning, where the so-called double descent (when more parameters are included than observations, the mean squared error goes down a second time) has put the traditional view on model selection upside down. Essentially, for adequate neighbourhood selection in models with a large number of parameters, the volume of the model space needs to be included in the penalty. Most neighbourhood selection methods (e.g., Lasso, AIC, BIC) lead to spurious edges (high false positive rate), but we prove that in the high-dimensional setting, minimum description length leads to correct neighbourhood selection or smaller (low false positive rates) in both cases when either the model is correctly or incorrectly assumed linear