发表机构
Vienna University of Technology(维也纳工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究神经网络训练最优性的计算复杂性,针对线性和ReLU激活函数网络,给出新算法上界,确定ReLU网络隐藏神经元出度为1时多项式时间可处理,找到线性激活函数网络满足新条件时首个多项式时间可解类。
AI 中文摘要
尽管神经网络在当代机器学习研究中具有基础性作用,但即使处理最简单的激活函数,我们对最优训练神经网络的计算复杂性理解仍不完整。近期虽有许多关于线性和ReLU激活函数下问题的更紧下界结果,但在识别新的多项式时间可处理网络架构方面进展较少。本文为训练线性和ReLU激活神经网络至最优性获得了新的算法上界,拓展了这些问题的可处理边界。特别是对于ReLU网络,确定了隐藏神经元出度为1的所有架构的多项式时间可处理性;对于线性激活函数网络,通过一种算法确定了首个非平凡的多项式时间可解网络类,该算法能最优训练满足新数据吞吐量条件的网络架构。
英文摘要
In spite of the fundamental role of neural networks in contemporary machine learning research, our understanding of the computational complexity of optimally training neural networks remains incomplete even when dealing with the simplest kinds of activation functions. Indeed, while there has been a number of very recent results that establish ever-tighter lower bounds for the problem under linear and ReLU activation functions, less progress has been made towards the identification of novel polynomial-time tractable network architectures. In this article we obtain novel algorithmic upper bounds for training linear- and ReLU-activated neural networks to optimality which push the boundaries of tractability for these problems beyond the previous state of the art. In particular, for ReLU networks we establish the polynomial-time tractability of all architectures where hidden neurons have an out-degree of $1$, improving upon the previous algorithm of Arora, Basu, Mianjy and Mukherjee. On the other hand, for networks with linear activation functions we identify the first non-trivial polynomial-time solvable class of networks by obtaining an algorithm that can optimally train network architectures satisfying a novel data throughput condition.
CommentsAppeared in the proceedings of NeurIPS 2023