AI 中文总结
该研究证明ReLU网络每增加一层可指数级减少神经元,得到相邻固定深度的首个指数层级,回答了相关学者的问题,还证明了更正则目标的深度精确分离。
AI 中文摘要
我们证明了ReLU神经网络的一种深度层级,其中每增加一个ReLU层都能指数级地减少神经元数量。对于所有$k\geq2$,我们构造了一个全局取值在$[0,1]$、1-Lipschitz函数,由宽度为$\mathcal{O}(d^4)$的深度-$(k+1)$网络实现;而在原点处呈指数距离支撑的绝对连续分布下,任何深度为$k$、权重无约束且宽度至多为$\frac{2^d}{2d(k-1)}$的网络,其平方$L_2$误差至少为$1/24$。据我们所知,这是首个在所有相邻固定深度间的指数层级,也是首个针对ReLU网络在两个固定深度间的指数分离,其中较浅网络的深度至少为3。该下界还直接得到了精确计算的对应层级。此外,当$k=2$时,得到了深度3与深度2间的紧支撑分离,且浅网络权重无约束,回答了Safran、Eldan和Shamir(2019)提出的问题。我们构造中使用的分布的所有质量都位于指数半径处,使该层级超出了正则性范围,在此范围内此类分离将意味着重大阈值电路下界。我们还证明了更正则目标的精确分离,该目标全局取值在$[0,1]$、$\mathcal{O}(\sqrt{d})$-Lipschitz,且将单位超立方体映射到$[0,1]$,它由多项式宽度的深度-4网络计算,而任何在单位超立方体上与它一致的深度-3网络,即使权重无约束,也需要指数级的第一层神经元。
英文摘要
We prove a depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons. For all $k\geq2$, we construct a globally $[0,1]$-valued, $1$-Lipschitz function realized by a depth-$(k+1)$ network of width $\mathcal{O}(d^4)$, whereas any depth-$k$ network with unrestricted weights and width at most $\frac{2^d}{2d(k-1)}$ has squared $L_2$ error at least $1/24$ under an absolutely continuous distribution supported at exponential distance from the origin. To the best of our knowledge, this is the first exponential hierarchy across all adjacent fixed depths, and the first exponential separation for ReLU networks between two fixed depths whose shallower network has depth at least $3$. The lower bound also immediately yields the corresponding hierarchy for exact computation. Moreover, the case $k=2$ gives a compactly supported separation between depths $3$ and $2$ with unrestricted shallow-network weights, answering a question raised by Safran, Eldan, and Shamir (2019). The distribution used in our construction nevertheless has all its mass at exponential radius, placing the hierarchy outside the regularity regime in which such a separation would imply major threshold-circuit lower bounds. We also prove an exact separation for a more regular target, which is globally $[0,1]$-valued and $\mathcal{O}(\sqrt d)$-Lipschitz and maps the unit hypercube onto $[0,1]$. It is computed by a polynomial-width depth-$4$ network, whereas any depth-$3$ network agreeing with it on the unit hypercube requires exponentially many first-layer neurons, even with unrestricted weights.