arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04999cs.LGstat.ML

超越过参数化:利用主动查询可证明地学习输入凸多层多项式网络

Beyond Overparameterization: Provable Learning of Input-Convex Multi-Layer Polynomial Networks with Active Queries

  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Jinqi Tang, Qian Chen, Shihong Ding, Cong Fang

AI总结:

针对过参数化局限,提出ASPIRE算法,利用输入凸性和主动查询,在多项式时间内以较低样本复杂度实现深层多项式网络的参数级恢复,首次证明高质量数据的指数级优势。

AI中文摘要:

多层神经网络的理论理解在很大程度上局限于过参数化设置,这掩盖了参数可辨识性并导致较高的样本复杂度。神经正切核(NTK)为宽网络提供了一般性理论,但并未提供高效的样本复杂度保证。最近的特征学习结果超越了针对单神经元、多索引和分层目标的核方法。然而,分析通常局限于浅层或特定架构以及过参数化机制。我们打破了这一范式,通过使用主动数据查询,实现了对深层目标网络的参数级恢复。具体而言,我们研究了具有偶数次$k$次单项式激活和非负高层权重的$L$层多项式网络。这种结构使目标网络输入凸,而优化景观相对于参数而言仍然高度非凸。利用输入凸性和主动查询,我们提出了\textbf{ASPIRE}(\textbf{A}ctive \textbf{S}am\textbf{P}ling for \textbf{I}terative \textbf{R}ecovery via \textbf{E}igendirections),一种基于逐层采样的对角化算法,该算法以多项式时间内的样本复杂度$\widetilde O_{k,L}\left(d^{L^2+O(L)}\delta^{-2e}\right)$将所有网络参数恢复到$\delta$精度。据我们所知,这是对指数仅随深度多项式增长的深层目标网络的第一个参数恢复保证,也是第一个证明在神经网络训练中使用高质量数据的有效性的论证,且具有显著的\emph{指数级}分离。

英文摘要:

The theoretical understanding of multi-layer neural networks is largely confined to overparameterized settings, which obscure parameter identifiability and incur high sample complexity. Neural tangent kernel (NTK) provides a general theory for wide networks, but does not offer efficient sample-complexity guarantees. Recent feature-learning results go beyond kernel methods for single-neuron, multi-index, and hierarchical targets. However, the analysis is often restricted to shallow or specific architectures and to the overparameterized regime. We break this paradigm to achieve parameter-level recovery of deep target networks, albeit by using active data queries. Specifically, we study $L$-layer polynomial networks with even degree-$k$ monomial activations and nonnegative higher-layer weights. This structure makes the target network input-convex, while the optimization landscape remains highly nonconvex with respect to the parameters. Leveraging input convexity and active queries, we propose \textbf{ASPIRE} (\textbf{A}ctive \textbf{S}am\textbf{P}ling for \textbf{I}terative \textbf{R}ecovery via \textbf{E}igendirections), a layerwise sampling-based diagonalization algorithm that recovers all network parameters to $δ$-accuracy with sample complexity $ \widetilde O_{k,L}\left(d^{L^2+O(L)}δ^{-2e}\right) $ in polynomial time. To our knowledge, this is the \emph{first} parameter-recovery guarantee for deep target networks whose exponent grows only polynomially with depth, as well as the \emph{first} justification for the effectiveness of using high-quality data in neural network training, with a remarkably \emph{exponential} separation.

↑