arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13749cs.LGstat.ML

作为 Grokking 极限状态的代数可表示性:具有全纯激活的精确可解模型

Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

Chon-Fai Kam, Xavier Cadet, Miloud Bessafi, Frederic Cadet

首次发表
浏览论文内容

中文总结 AI 辅助

研究在模运算训练的神经网络,当可表达函数类坍缩到有限维代数簇时的情况。通过全纯单项式激活的两层网络及代数表征研究,给出训练损失下界,实验表明代数预测准确率高,呈现二元行为,是容量 - Grokking关系极限,瓶颈消融连接极端与标准网络。

中文摘要 AI 辅助

在模运算上训练的神经网络表现出 Grokking,即从记忆到泛化的延迟转变,这取决于模型容量。当架构的可表达函数类坍缩到有限维代数簇时会发生什么?我们研究了具有全纯单项式激活σ(z)=z^k的两层网络,在通过单位根编码的模任务上进行训练。网络输出限于(Z_p)^2特征的(k + 1)维子空间。我们给出了该子空间的完整代数表征:当且仅当离散傅里叶支持位于对角线u + v = k (mod p)上时任务可表示。这不仅限制最终泛化,也限制记忆本身。不可表示的目标即使在训练集上也无法拟合,我们证明了训练损失的正下界。在585次运行中,代数预测与观察结果的准确率为99.8%,无记忆阶段和Grokking,结果分为即时成功和彻底失败。这种二元行为是容量 - Grokking关系的极限情况。瓶颈消融将此极端情况与标准网络联系起来。

英文摘要

Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately. What happens at the extreme of this spectrum, when the architecture's expressible function class collapses to a finite-dimensional algebraic variety? We study two-layer networks with a holomorphic monomial activation sigma(z)=z^k, trained on modular tasks encoded via roots of unity. Here the network output, regardless of hidden width, is confined to a (k+1)-dimensional subspace of characters of (Z_p)^2, an O(k/p^2) slice of the full function space. We give a complete algebraic characterisation of this subspace: a task is representable if and only if its discrete Fourier support lies on the diagonal u+v = k (mod p), which for linear-phase targets reduces to the arithmetic criterion m+n=k. This is not merely a constraint on eventual generalisation but on memorisation itself: because the outputs are algebraically confined, a non-representable target cannot be fit even on the training set, and we prove a positive lower bound on the training loss, independent of width. Across 585 runs the algebraic prediction matches the observed outcome with 99.8% accuracy, with no memorisation regime and no grokking; outcomes split cleanly into instant success and outright failure. This binary behaviour is the limiting case of the capacity-grokking relationship: when the expressible class shrinks to a fixed algebraic object, the question of when a network will grok dissolves into whether it can represent the target at all. A bottleneck ablation connects this extreme to standard networks, tracing a continuous path from representational failure, through memorisation without generalisation, to grokking with a shrinking gap as capacity grows.

发表机构

  • University Paris City & University of Reunion(巴黎城市大学和留尼汪大学)
  • Dipartimento di Fisica e Chimica Emilio Segrè, Università degli Studi di Palermo(巴勒莫大学埃米利奥·塞格雷物理与化学系)
  • Thayer School of Engineering, Dartmouth College(达特茅斯学院塞耶工程学院)
  • EnergyLab, University of Reunion(留尼汪大学能源实验室)
  • PEACCEL, AI for Biologics(PEACCEL,生物制品人工智能)

机构由 AI 辅助整理,请以论文原文为准。

↑