无维度随机超平面排列的秩提升
Dimension-Free Rank Lifting from Random Hyperplane Arrangements
查看机构详情
- Sapienza University of Rome(罗马第一大学)
- EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文研究神经网络随机隐藏层实现秩提升所需的宽度,证明正齐次非多项式激活函数在无维度界下可实现精确秩提升,并统一推广了稳定秩提升的保证。
中文摘要 AI 辅助
我们研究了神经网络随机初始化的隐藏层实现秩提升所需的宽度。具体而言,给定一个数据集 $X \in \mathbb{R}^{m \times d}$,包含 $m$ 个 $d$ 维输入向量,且任意两个向量之间的夹角至少为 $\theta$,我们考虑随机特征矩阵 $\sigma(XR)$,其中 $R$ 为标准高斯矩阵。对于正齐次非多项式激活函数(包括符号函数、Heaviside 函数、ReLU 及其幂函数等),我们证明当 $$n \gtrsim \frac{1}{\theta}\max\left\{m,\log\left(\frac{1}{\delta}\right)\right\}$$ 个神经元时,$\sigma(XR)$ 以至少 $1-\delta$ 的概率具有满行秩 $m$。这一无维度界在指数级别上改进了先前针对符号特征的通用维度保证(Drago 等人,2026),并且本质上是紧的。证明表明,一个随机特征列以 $\Omega(\theta)$ 的概率逃离 $\mathbb{R}^m$ 的每个真子空间,这利用了邻近高斯方向的耦合以及诱导超平面排列的局部穿越。我们还研究了稳定秩提升,其目标是建立精确秩提升的定量类比,即在高概率下对经验特征 Gram 矩阵的最小特征值给出下界。我们的分析统一并推广了所有 $q$-齐次非多项式激活函数的稳定秩保证,这些保证先前见于 Panigrahi 等人(2020)和 Song(2026)的工作。特别地,我们将总体核的对角占优 Taylor 尾部与截断和矩阵集中相结合,证明对于正齐次非多项式激活函数,稳定秩提升在宽度 $$n \gtrsim C^q \frac{m}{\theta^{2q+1}} \log^{2q+\frac{1}{2}}\left(\frac{m}{\theta}\right) \log\left(\frac{m}{\delta}\right)$$ 时实现,其中 $q$ 是激活函数的次数,$C > 0$ 是某个通用常数。
英文摘要
We study the width required for a randomly initialized hidden layer of a neural network to achieve rank lifting. Namely, given a dataset $X \in \mathbb{R}^{m \times d}$ of $m$, $d$-dimensional input vectors separated by an angle of at least $θ$, we consider the random feature matrix $σ(XR)$, where $R$ is standard Gaussian. For positively homogeneous nonpolynomial activations, which include sign, Heaviside, ReLU, and ReLU powers among others, we prove that $$n \gtrsim \frac{1}θ\max\left\{m,\log\left(\frac{1}δ\right)\right\}$$ neurons suffice for $σ(XR)$ to have full row rank $m$ with probability at least $1-δ$. This dimension-free bound exponentially improves the previous general-dimensional guarantee for sign features (Drago et al., 2026) and is essentially tight. The proof shows that one random feature column escapes every proper subspace of $\mathbb{R}^m$ with probability $Ω(θ)$, using a coupling of nearby Gaussian directions and a local crossing of the induced hyperplane arrangement. We also study stable rank lifting, where the goal is to establish a quantitative analogue of exact rank lifting, i.e., a lower bound on the smallest eigenvalue of the empirical feature Gram matrix in high-probability. Our analysis unifies and generalizes stable rank guarantees for all $q$-homogeneous non-polynomial activations following prior work in Panigrahi et al. (2020) and Song (2026). In particular, we combine a diagonally dominant Taylor tail of the population kernel with truncation and matrix concentration, to show that for positively homogeneous nonpolynomial activations, stable rank lifting is achieved at width $$n \gtrsim C^q \frac{m}{θ^{2q+1}} \log^{2q+\frac{1}{2}}\left(\frac{m}θ\right) \log\left(\frac{m}δ\right),$$ where $q$ is the degree of the activation and $C > 0$ is some universal constant.