arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

浅层ReLU网络中的总体损失:偏置与临界点族

Population loss in shallow ReLU networks: Bias & families of critical points

Michael Field

arXiv 2609.30661首次发表:更新:

发表机构

UC Santa Barbara(加州大学圣塔芭芭拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文推导了带偏置浅层ReLU网络在学生-教师核模型中的总体损失公式,证明添加偏置严格降低损失且伪极小值族可扩展,并利用Owen T函数与解析几何工具分析景观几何。

AI 中文摘要

本文的主要结果是给出了学生-教师核模型中总体损失的一个公式,该公式适用于带偏置的浅层ReLU网络。这扩展了Choo和Saul(2009)以及Brutzkus和Globerson(2017)的先前工作。该公式本质上利用了Owen的T函数。文中给出了T函数的必要理论,并基于Komelj(2023)的算法,使用MPFR对T函数进行了高精度编码,该编码可按需提供。研究表明,Arjevani和作者过去论文中描述的各种伪极小值族可扩展到带偏置的网络,并且添加偏置后损失总是严格减少。由添加偏置引起的景观几何变化似乎相对温和。本文仅描述了最简单的例子,其中假设输入数量等于神经元数量(此限制出于篇幅原因)。文中还回顾了先前关于无偏网络的相关结果。除高斯统计外,主要的数学工具和思想来自解析几何(解析集和子解析集、曲线选择引理)。

英文摘要

The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families of spurious minima described in past papers of Arjevani and the author extend to biased networks and that the loss is always strictly decreased when bias is added. The change in landscape geometry caused by adding bias appears to be relatively mild. Only the simplest examples are described in this paper where it is assumed that the number of inputs is equal to the number of neurons (this restriction is for reasons of length). A review of relevant previous results on unbiased networks is included. Aside from Gaussian statistics, the main mathematical tools and ideas come from analytic geometry (analytic and subanalytic sets, the Curve Selection Lemma).

Comments118 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑