arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稀疏激活神经网络近紧的 Rademacher 界

Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks

Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang, Junbin Gao, Andi Han

arXiv 2609.09130首次发表:更新:

发表机构

University of New South Wales; University of Texas at Austin; University of Sydney(新南威尔士大学; 德克萨斯大学奥斯汀分校; 悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究单隐层 ReLU 网络中输入依赖稀疏性的统计复杂度,提出近紧的 Rademacher 界,并证明宽度依赖性与输入域的关键影响。

AI 中文摘要

一个输入可能激活少数隐藏单元,即使不同输入共同使用整个网络。我们研究了 Awasthi 等人 (COLT 2024) 的单隐层 ReLU 模型中这种输入依赖稀疏性的统计复杂度。对于宽度 $s$、每个输入至多 $k$ 个活跃单元、有效权重和偏置界 $W,B$,该类固定半径 $R$ 输入域中的每个大小为 $m$ 的样本满足 $\mathcal{R}(S)\le CWR\min\{k,\sqrt{sk/m}\log^{3/2}(2m)\}+kB/\sqrt m$。一个保持支撑的覆盖和一个归一化的链式论证去除了之前的显式维度因子,直到对数因子。适当 i.i.d. 边际分布上的下界与这些对数因子匹配,表明跨输入改变活跃单元保留了宽度依赖性。输入域很重要:在整个球上稀疏的零偏置网络至多有 $2k$ 个非零单元,复杂度为 $O(kWR/\sqrt m)$,而偏置界与 $WR$ 相当时,在同一域上仅在对数维度中恢复最坏情况速率。一个球冠构造证明了后一主张,而无需假设仅在采样支撑上的稀疏性。对于指定的归一化有界损失和与 $WR$ 相当的偏置,我们还获得了阶为 $\min\{1,\sqrt{s/(km)}\}$(直到对数因子)的不可知极小极大超额风险界。

英文摘要

An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width $s$, at most $k$ active units per input, and effective weight and bias bounds $W,B$, every size-$m$ sample in the class's fixed radius-$R$ input domain satisfies $\mathcal{R}(S)\le CWR\min\{k,\sqrt{sk/m}\log^{3/2}(2m)\}+kB/\sqrt m$. A support-preserving cover and a single normalized chaining argument remove the previous explicit dimension factor, up to logarithms. Lower bounds on appropriate i.i.d. marginals match up to those logarithms, showing how changing active units across inputs retains a width dependence. The input domain matters: zero-bias networks sparse on the entire ball have at most $2k$ nonzero units and complexity $O(kWR/\sqrt m)$, whereas bias bounds comparable to $WR$ restore the worst-case rate on that same domain in only logarithmic dimension. A spherical-cap construction proves the latter claim without assuming sparsity merely on the sampling support. For a specified normalized bounded loss and biases comparable to $WR$, we also obtain agnostic minimax excess-risk bounds of order $\min\{1,\sqrt{s/(km)}\}$ up to logarithms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑