发表机构
Columbia University; Anthropic; UT Austin(哥伦比亚大学; Anthropic; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种算法,在在线学习模型下以$2^{n-c_\u03b3 n}$时间学习大小为$n^\u03b3$的深度为3的电路,实现多项式节省,通过改进CNF近似度和随机游走谱放大。
AI 中文摘要
我们研究了在可实现在线学习的错误界模型下学习深度为3的电路的挑战性问题,该模型比分布无关的PAC学习更困难。先前针对此问题的算法(由Servedio和Tan [ST17]提出)只能学习大小为poly$(n)$的深度为3的电路,其运行时间为$2^{n - \frac{n}{\text{log} n}}$,因此其运行时间为$N^{1-o(1)}$,其中$N=2^n$是朴素记忆化方法的运行时间。在本工作中,我们大幅改进了[ST17]的结果:对于任意常数$\u2265 1$,我们给出一个算法,学习大小为$n^\u03b3$的深度为3的电路,运行时间为$2^{n-c_\u03b3 n}$,其中$c_\u03b3>0$仅依赖于$\u03b3$而不依赖于$n$。因此,我们在学习任意多项式大小的深度为3的电路时,相对于朴素方法实现了多项式节省。我们改进的主要驱动力是对宽度为$k$的CNF的近似度的改进界。受Szegedy [Sze04]和Magniez等人 [MNRS11]的启发,我们构造的粗略思路是使用切比雪夫多项式来有效放大精心设计的随机游走的谱间隙。这结合了类似随机限制的方法,分别学习对应于随机选择的变量集的不同赋值的不同子函数,使用在特殊设计的特征空间上的感知机算法。我们方法的简化预热实例实现了$c_\u03b3 = \text{exp}(-O(\u03b3))$;通过向该预热添加进一步成分,我们获得了结果的尖锐形式,实现了$c_\u03b3=\u03a9(1)/\u03b3$。
英文摘要
We study the challenging problem of learning depth-three circuits in the mistake-bound model of (realizable) online learning, which is a more difficult model than distribution-free PAC learning. Prior algorithms for this problem, due to Servedio and Tan [ST17], could only learn polynomial-size depth-three circuits of poly$(n)$ size over $\{0,1\}^n$ with a running time of $2^{n - Ω(n/\log n)}$, and hence they ran in time $N^{1-o(1)}$ where $N=2^n$ is the running time of a naive memorization-based approach. In this work we substantially improve on the [ST17] result: for any constant $γ\geq1$, we give an algorithm that learns depth-three circuits of size $n^γ$ with running time \[ 2^{n-c_γn}, \] where $c_γ>0$ depends only on $γ$ and not on $n$. Hence we achieve a polynomial savings over the naive approach for learning any polynomial-size depth-three circuit. The main driving force behind our improvement is an improved bound on the approximate degree of width-$k$ CNFs. Inspired by Szegedy [Sze04] and Magniez et al. [MNRS11], the rough idea of our construction is to use a Chebyshev polynomial to efficiently amplify the spectral gap of a carefully designed random walk. This is combined with a random-restriction-like approach to separately learn different subfunctions corresponding to different assignments to a randomly chosen set of variables, using the Perceptron algorithm over a specially designed feature space. A simplified warmup instantiation of our approach achieves $c_γ= \exp(-O(γ))$; by augmenting this warmup with further ingredients we obtain the sharp form of our result, which achieves $c_γ=Ω(1)/γ$.