arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

常数深度和对数深度神经网络之间的算法分离

Algorithmic Separation between Constant-Depth and Logarithmic-Depth Neural Networks

Yunwei Ren, Zihao Wang, Jason D. Lee

arXiv 2607.25200首次发表:更新:

发表机构

Princeton University; Stanford University; University of California, Berkeley(普林斯顿大学; 斯坦福大学; 加利福尼亚大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究常数深度和对数深度神经网络的算法分离,识别出一类布尔函数,对数深度网络能用逐层坐标下降法通过重构谱高效学习,还展示了一个子类,常数深度网络在此子类上有常数\(L^2\)逼近误差。

AI 中文摘要

尽管深度网络在经验上比浅层网络有优势,但理论深度分离主要涉及逼近能力,算法结果大多限于二三层网络比较。本文证明了常数深度和对数深度网络之间的首个算法分离。具体而言,识别出一类具有分层结构傅里叶谱的布尔函数,对数深度网络可通过分层和自适应重构谱,用逐层坐标下降法高效学习。还展示了一个子类,对于该子类,在超立方体上的均匀分布下,每个具有足够正则激活和受控谱范数的常数深度、多项式宽度网络必定会产生常数\(L^2\)逼近误差。

英文摘要

Despite the empirical advantages of deep networks over shallow ones, theoretical depth separations largely concern approximation power, while algorithmic results are mostly limited to comparisons between two- and three-layer networks. In this work, we prove the first algorithmic separation between constant-depth and logarithmic-depth networks. Specifically, we identify a class of Boolean functions with hierarchically structured Fourier spectra that logarithmic-depth networks can learn efficiently using layerwise coordinate descent by reconstructing the spectra hierarchically and adaptively. We also exhibit a subclass for which every constant-depth, polynomial-width network with sufficiently regular activations and controlled spectral norms must incur constant $L^2$ approximation error under the uniform distribution over the hypercube.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑