发表机构
UC Santa Barbara(加州大学圣塔芭芭拉分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文推导了带偏置浅层ReLU网络在学生-教师核模型中的总体损失公式,证明添加偏置严格降低损失且伪极小值族可扩展,并利用Owen T函数与解析几何工具分析景观几何。
AI 中文摘要
本文的主要结果是给出了学生-教师核模型中总体损失的一个公式,该公式适用于带偏置的浅层ReLU网络。这扩展了Choo和Saul(2009)以及Brutzkus和Globerson(2017)的先前工作。该公式本质上利用了Owen的T函数。文中给出了T函数的必要理论,并基于Komelj(2023)的算法,使用MPFR对T函数进行了高精度编码,该编码可按需提供。研究表明,Arjevani和作者过去论文中描述的各种伪极小值族可扩展到带偏置的网络,并且添加偏置后损失总是严格减少。由添加偏置引起的景观几何变化似乎相对温和。本文仅描述了最简单的例子,其中假设输入数量等于神经元数量(此限制出于篇幅原因)。文中还回顾了先前关于无偏网络的相关结果。除高斯统计外,主要的数学工具和思想来自解析几何(解析集和子解析集、曲线选择引理)。
英文摘要
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families of spurious minima described in past papers of Arjevani and the author extend to biased networks and that the loss is always strictly decreased when bias is added. The change in landscape geometry caused by adding bias appears to be relatively mild. Only the simplest examples are described in this paper where it is assumed that the number of inputs is equal to the number of neurons (this restriction is for reasons of length). A review of relevant previous results on unbiased networks is included. Aside from Gaussian statistics, the main mathematical tools and ideas come from analytic geometry (analytic and subanalytic sets, the Curve Selection Lemma).
Comments118 pages, 3 figures