发表机构
Anhui iFlytek Yinglian Technology Co., Ltd.(安徽讯飞英联科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Topo²框架,将深度网络的记忆与泛化在因果上分离,建立相关规律,验证其特性,将记忆转化为可测量的拓扑层,为二者的研究提供了可量化的工具。
AI 中文摘要
在带噪声标签上训练的深度网络,会同时在干净数据上实现泛化,并对翻转的标签进行记忆,这两种现象通常被混同为对单一能力的压力。我们提出了Topo²,这一测量框架可将二者在因果上分离、可测量且符合规律。表征空间的持续同调H1结构,可分离为类内流形通道(取决于训练停止点)和类间通道(对已记忆的翻转样本的单调读出)。干预措施FM0处方(第0轮对翻转样本的零损失)可达到每种设置的泛化上限,同时几乎不进行记忆。在该框架内,我们建立了一套具有分级证据的规律集合:(L2)FM0分离处方(9/9);(L1)类内通道作为训练位置函数(中上升阶段6/6;收敛回溯CIFAR 3/3、SVHN 2/3);(L3)环构造恒等式(定义性的,非规律);以及TLS(记忆-泛化拓扑分层):记忆具有因果可加性、锚定性(使干净样本沉默会破坏表征)、可逆性(剥离记忆可恢复接近上限的泛化),且可量化计费(记忆成本规律,参考能力下有效斜率系数C约为0.38:CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715,通常依赖于能力,可追溯至干净样本特征位移)。我们还公布了该框架的边界:一份包含9个死胡同的证伪台账,以及一份工具验证部分,排除了6类全局统计量作为类内通道的解释。该框架将“记忆”从一种定义模糊的能力转变为可测量、可分离、可逆的拓扑层。
英文摘要
Deep networks trained on noisy labels simultaneously generalize on clean data and memorize flipped labels. These are usually conflated as pressures on one capacity. We present Topo^2, a measurement framework that makes them causally separable, measurable, and law-governed. Persistent-homology H1 structure of the representation space separates into a within-class manifold channel (a function of the training stopping point) and a cross-class channel (a monotone readout of memorized flipped samples). An intervention, the FM0 prescription (zero loss on flipped samples from epoch 0), reaches each setting's generalization ceiling while memorizing essentially nothing. Within the framework we establish a law set with graded evidence: (L2) FM0 separation prescription (9/9); (L1) the within-channel as a training-position function (mid-rise 6/6; convergence-back CIFAR 3/3, SVHN 2/3); (L3) a ring-construction identity (definitional, not a law); and TLS (memory-generalization topological layering): memory is causally additive, anchored (silencing clean collapses the representation), invertible (stripping memory restores near-ceiling generalization), and quantitatively billable (the memorization cost law, effective slope coefficient C ~ 0.38 at the reference capacity: CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715, capacity-dependent in general and traced to clean-sample feature displacement). We also publish the framework's boundaries: a falsification ledger of nine dead ends, and an instrument-vindication section that excludes six families of global statistics as explanations of the within-channel. The framework turns "memorization" from an ill-defined capacity into a measurable, separable, invertible topological layer.
Comments14 pages, 6 figures