arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25221cs.LGmath.OCstat.ML

用双隐藏层ReLU网络表示MAX函数

Representing MAX functions using two-hidden-layer ReLU networks

Zhimao Wang, Amitabh Basu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对MAX_N函数的双隐藏层ReLU网络表示问题,通过计算机辅助搜索独立得到MAX₅至MAX₈的表示,与近期Ruess等人的成果存在细微差异,为相关研究提供了新进展。

中文摘要 AI 辅助

我们研究使用双隐藏层ReLU神经网络对MAX_N(x)=max{x₁,…,x_N}的精确表示。近年来,为了刻画表示连续分段线性函数所需隐藏层的精确数量,该问题已得到研究。当前最佳下界为2,而上界为关于N的对数函数。至今仍完全未知正确答案是否为常数数量的隐藏层(甚至可能仅为2层)。事实上,近期的一项突破是[Bakaev等人,2026]将MAX₅表示为双隐藏层ReLU函数,该论文指出N≥6时的MAX_N问题仍未解决。通过精心设计的计算机辅助搜索,我们得到了MAX₅、MAX₆、MAX₇和MAX₈的双隐藏层ReLU表示。我们通过考虑形如max{∑_{r=1}^s max(x_{a_r},x_{b_r}), ∑_{r=1}^s max(x_{c_r},x_{d_r})}项的有理线性组合来获得这些表示,其中a_r、b_r、c_r、d_r∈{1,…,N}。每个坐标对的内部最大值可在第一隐藏层计算,两侧和的外部最大值可在第二隐藏层计算。因此,这些项的每个有限线性组合都有一个双隐藏层ReLU实现。因此,这种形式的MAX_N恒等式给出了MAX_N的精确双隐藏层ReLU表示。就在近期,[Ruess等人,2026]获得了所有N≤10的上述形式的MAX_N双隐藏层表示。我们的表示不同且是独立开发的。尽管我们的技术与[Ruess等人,2026]提出的大部分高层思想一致,但也存在一些可能对该问题未来研究有意义的细微差异。

英文摘要

We study exact representations of $\mathrm{MAX}_N(x)=\max{x_1,\ldots,x_N}$ using two-hidden-layer ReLU neural networks. This problem has been studied in recent years in an attempt to characterize the exact number of hidden layers required to represent continuous piecewise linear functions. The best lower bound is 2, while the current upper bound is logarithmic in $N$. It remains completely open if the right answer is a constant number of hidden layers (possibly even 2!) or not. In fact, a recent breakthrough was the representation of $\mathrm{MAX}_5$ as a two-hidden-layer ReLU function obtained in [Bakaev et al., 2026], and the case of $\mathrm{MAX}_N$ was stated as open for $N\geq 6$ in that paper. Using a careful computer assisted search, we obtain two-hidden-layer ReLU representations of $\mathrm{MAX}_5, \mathrm{MAX}_6, \mathrm{MAX}_7$, and $\mathrm{MAX}_8$. We obtain these by considering rational linear combinations of terms of the form $\max\{\sum_{r=1}^{s}\max(x_{a_r},x_{b_r}),\sum_{r=1}^{s}\max(x_{c_r},x_{d_r})\}$, where $a_r,b_r,c_r,d_r\in\{1,\ldots,N\}$. Each inner maximum of two coordinates can be computed in a first hidden layer, and the outer maximum of the two side-sums can be computed in a second hidden layer. Consequently, every finite linear combination of these terms has a two-hidden-layer ReLU realization. An identity for $\mathrm{MAX}_N$ in this form therefore gives an exact two-hidden-layer ReLU representation of $\mathrm{MAX}_N$. Very recently, two-hidden-layer representations of $\mathrm{MAX}_N$ of the above form were obtained for all $N\leq 10$ in [Ruess et al., 2026]. Our representations are different and were developed independently. While our techniques share most of the high-level ideas presented in [Ruess et al., 2026], there are also some minor differences which may be of interest for future research on this problem.

发表机构

  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑