arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

最小范数单变量两层ReLU分类:精确解与带跳跃连接的全局最优性

Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections

Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazić

arXiv 2609.28438首次发表:更新:

发表机构

University of Warsaw; University of Warwick; INRIA & LMO, Université Paris-Saclay(华沙大学; 华威大学; 法国国家信息与自动化研究所与巴黎萨克雷大学LMO)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究单变量两层ReLU网络的最小范数分类,刻画了有无偏置惩罚时的精确解,并证明添加跳跃连接使所有KKT点全局最优,同时揭示了稀疏性限制。

AI 中文摘要

我们研究了单变量两层ReLU网络用于二分类时的最小范数插值和ℓ2正则化逻辑损失最小化问题。我们在函数空间中给出了最优分类器的完整几何刻画,解决了解如何依赖于参数范数中是否包含隐层偏置的问题。当偏置不受惩罚时,最小范数插值器恰好是连续分段仿射函数,它们紧贴每个标签切换点,并在适当方向具有拐点。当偏置受惩罚时,函数空间中的最小化器是唯一的,在每个中间同标签段中恰好有一个拐点,因此是最稀疏的正间隔分类器。我们进一步证明,添加一个自由的仿射跳跃连接不会改变这些函数空间解,但从根本上改善了参数空间的地形:约束问题的每个KKT点都成为全局最优,而没有跳跃连接时可能出现次优的KKT点。对于足够弱的ℓ2正则化逻辑损失,我们建立了类似的全局最优性和几何结果。在偏置不受惩罚的情况下,我们识别出一个额外的类稀疏性限制,这意味着大多数最小范数插值器不能作为间隔归一化逻辑损失最小化器的小正则化极限出现。跨不同数据集复杂度和网络宽度的数值实验支持了预测的地形和稀疏性现象。

英文摘要

We study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-layer biases are included in the parameter norm. When biases are unpenalized, the minimal-norm interpolators are exactly the continuous piecewise-affine functions that hug every label switch and have kinks of the appropriate convexity. When biases are penalized, the minimizer is unique in function space, has exactly one kink in each intermediate same-label segment, and is therefore a sparsest positive-margin classifier. We further show that adding a free affine skip connection leaves these function-space solutions unchanged but fundamentally improves the parameter-space landscape: every KKT point of the constrained problem becomes globally optimal, whereas suboptimal KKT points can occur without the skip connection. We establish analogous global-optimality and geometric results for sufficiently weak $\ell_2$-regularization of the logistic loss. In the unpenalized-bias case, we identify an additional sparsity-like restriction, implying that most minimal-norm interpolators cannot arise as small-regularization limits of margin-normalized logistic-loss minimizers. Numerical experiments across varying dataset complexity and network width support the predicted landscape and sparsity phenomena.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑