发表机构
Princeton University; Stanford University; University of California, Berkeley(普林斯顿大学; 斯坦福大学; 加利福尼亚大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究常数深度和对数深度神经网络的算法分离,识别出一类布尔函数,对数深度网络能用逐层坐标下降法通过重构谱高效学习,还展示了一个子类,常数深度网络在此子类上有常数\(L^2\)逼近误差。
AI 中文摘要
尽管深度网络在经验上比浅层网络有优势,但理论深度分离主要涉及逼近能力,算法结果大多限于二三层网络比较。本文证明了常数深度和对数深度网络之间的首个算法分离。具体而言,识别出一类具有分层结构傅里叶谱的布尔函数,对数深度网络可通过分层和自适应重构谱,用逐层坐标下降法高效学习。还展示了一个子类,对于该子类,在超立方体上的均匀分布下,每个具有足够正则激活和受控谱范数的常数深度、多项式宽度网络必定会产生常数\(L^2\)逼近误差。
英文摘要
Despite the empirical advantages of deep networks over shallow ones, theoretical depth separations largely concern approximation power, while algorithmic results are mostly limited to comparisons between two- and three-layer networks. In this work, we prove the first algorithmic separation between constant-depth and logarithmic-depth networks. Specifically, we identify a class of Boolean functions with hierarchically structured Fourier spectra that logarithmic-depth networks can learn efficiently using layerwise coordinate descent by reconstructing the spectra hierarchically and adaptively. We also exhibit a subclass for which every constant-depth, polynomial-width network with sufficiently regular activations and controlled spectral norms must incur constant $L^2$ approximation error under the uniform distribution over the hypercube.