arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08503cond-mat.stat-mechcond-mat.dis-nncond-mat.str-el

副本碎裂与奇偶校验学习中的玻璃态动力学

Replica Fragmentation and Glassy Dynamics in Parity Learning

发表机构法国国家科学研究中心,巴黎综合理工学院,巴黎理工学院
查看机构详情
  • CPHT, CNRS, École Polytechnique, Institut Polytechnique de Paris(法国国家科学研究中心,巴黎综合理工学院,巴黎理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Han Ma

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过副本分析揭示Transformer在奇偶校验学习中的三种状态(记忆化、退缩、恢复),并利用自-交叉间隙及残差相关性区分持续记忆化与退缩-恢复动力学。

中文摘要 AI 辅助

我们研究了独立训练的Transformer神经网络如何从其局部畴壁重建二进制字符串。共享数据和训练协议的运行可以实现不同的函数。我们将它们视为副本,并测量真值对齐$m$、预测置信度$q_{\mathrm{self}}$和跨副本一致性$q_{\mathrm{cross}}$。自信的不一致定义了有限尺寸的副本碎裂,我们称之为类玻璃态。在小训练集下,副本能正确预测所有训练示例,但对未见输入仍保持高置信度的错误预测,我们将此状态称为记忆化。在较大训练集下,运行能够泛化,随后发生退缩。退缩发生在输出开始偏离真值而置信度仍保持较高时。已学习与未学习输出之间的前沿向较短字符串退缩。许多副本随后随着前沿再次推进而恢复。自-交叉间隙$q_{\mathrm{self}}-q_{\mathrm{cross}}$清晰地区分了记忆化、退缩和恢复三种学习状态。在扣除整体和位置相关的真值对齐后,它们的残差相关性也不同:记忆化副本具有弱且均匀相关的残差,而退缩状态具有最大比例的副本对,其残差呈反相关。也就是说,在某一输入上,若一对副本中的一个表现优于其平均值,则另一个往往表现较差。这些有限尺寸观察区分了持续记忆化与持续的退缩-恢复动力学。热力学极限和长时间极限,以及向其他学习任务的扩展,仍有待研究。

英文摘要

We study how independently trained Transformer neural networks reconstruct a binary string from its local domain walls. Runs sharing the data and training protocol can realize different functions. We treat them as replicas and measure truth alignment $m$, prediction confidence $q_{\mathrm{self}}$, and cross-replica agreement $q_{\mathrm{cross}}$. Confident disagreement defines the finite-size replica fragmentation that we call glass-like. With small training sets, replicas predict all training examples correctly but remain confident in incorrect predictions for unseen inputs, a regime we call memorization. With larger sets, runs can generalize and then retreat. Retreat occurs when outputs start to deviate from truth while confidence remains high. The frontier between learned and unlearned outputs recedes toward shorter strings. Many later recover as the frontier advances again. The self--cross gap $q_{\mathrm{self}}-q_{\mathrm{cross}}$ clearly distinguishes the three learning regimes of memorization, retreat, and recovery. With overall and position-dependent truth alignment subtracted, their residual correlations also differ: memorizing replicas have weakly and uniformly correlated residuals, while retreat has the largest fraction of replica pairs whose residuals are anti-correlated. That is, on the inputs where one replica of such a pair does better than its average, the other tends to do worse. These finite-size observations distinguish persistent memorization from ongoing retreat--recovery dynamics. The thermodynamic and long-time limits, and extensions to other learning tasks, remain open.

补充信息

↑