arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13057cs.LGstat.ML

良性损失景观可以与最坏情况下的困难共存

Benign Loss Landscapes Can Coexist with Worst-Case Hardness

Zach Furman, Stephan Wäldchen, Yangda Bei, Liam Hodgkinson

首次发表
浏览论文内容

中文总结 AI 辅助

该研究证明树张量网络虽包含难以学习的目标,但其损失景观对可实现目标是良性的,困难源于高阶退化鞍点,为理解深度学习的泛化提供了新视角。

中文摘要 AI 辅助

深度神经网络具有足够的表达能力,能够包含可在多项式时间内评估但无法通过梯度下降在多项式时间内学习的最坏情况目标。然而,对于实际任务,它们仍然学习得很好,这引发了问题:真实世界目标的何种非通用结构使得这种学习成为可能。现有的替代模型无法提出这个问题,因为它们要么完全缺乏难以学习的目标(如深度线性网络),要么无法有效评估此类目标(如核方法、无限宽度极限)。我们研究了树张量网络(TTNs),这是一个概括了深度线性网络和Tucker分解的模型类别。我们证明它们可以嵌入任意的只读布尔公式,因此包含多项式大小的目标,这些目标在神经网络相同的机制下无法通过梯度下降在多项式时间内学习。尽管如此,我们证明了对于每个可实现的目标,其损失景观是条件良性的:每个最小范数的局部最小值都是全局最小值。因此,令人惊讶的是,坏的局部最小值并不是区分TTNs中典型问题和最坏情况问题的因素。相反,TTNs中的学习困难可能源于高阶退化鞍点,我们证明这些鞍点是由秩亏导致的。通过对奇偶校验函数的案例研究进行了探讨,说明了TTNs将景观几何与计算困难联系起来的潜力。

英文摘要

Deep neural networks are expressive enough to contain worst-case targets that can be evaluated in polynomial time but cannot be learned in polynomial time by gradient descent. For practical tasks they nonetheless learn well, raising the question of what non-generic structure of real-world targets enables this. Existing surrogate models cannot pose this question because they either lack hard-to-learn targets entirely (deep linear networks) or cannot evaluate such targets efficiently (kernel methods, infinite-width limits). We study tree tensor networks (TTNs), a model class that generalizes deep linear networks and Tucker decompositions. We show they embed arbitrary read-once Boolean formulas, and thus contain polynomial-size targets that cannot be learned by gradient descent in polynomial time under the same mechanism as neural networks. Despite this, we prove that their loss landscapes are conditionally benign for every realizable target: every local minimum that is minimum-norm is global. Thus, surprisingly, bad local minima are not what distinguishes between typical and worst-case problems in TTNs. Instead, learning difficulty in TTNs can arise from high-order degenerate saddle points, which we show are caused by rank-deficiency. This is explored through a case study of the parity function, illustrating the potential for TTNs to relate landscape geometry to computational hardness.

发表机构

  • University of Melbourne(墨尔本大学)
  • Iliad

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑