arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

善度衡量什么?前馈学习的似然比解释

Gauge-Fixing the Forward-Forward Objective: A Whitened Goodness Derived from a Likelihood-Ratio Account

Paolo Giannitrapani

arXiv 2607.12501首次发表:更新:

发表机构

Sapienza University of Rome(罗马第一大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究前馈学习中善度衡量问题,基于显式生成模型指出平方善度是似然比检验充分统计量,FF阈值是其边界,还阐述了该方法在不同总体及层间归一化方面的特性及优势。

AI 中文摘要

前馈(FF)算法在局部训练每一层,使得一个标量善度(激活平方和)在真实输入上较高,在对比输入上较低,层间激活进行归一化。这些选择通常被视为启发式方法。在显式生成模型下并非如此:平方善度是两个尺度不同的零均值总体之间似然比检验的充分统计量,FF阈值是其边界。它具有一般性:各向异性总体产生马氏善度,普通平方是其各向同性情况;重尾总体产生饱和统计量,其斜率是后验精度——除法归一化——具有有界证据且仅在聚合下有优势。相同视角刻画了层间归一化:它必须消除长度同时保留每个坐标的能量,解释了我们在单位范数归一化下观察到的深度坍缩;成对目标允许一个尺度膨胀捷径,白化善度可消除该捷径。

英文摘要

The Forward-Forward algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and low on contrastive ones. Under an explicit generative model this goodness is the sufficient statistic of a likelihood-ratio test, and the pairwise form of the objective admits a gauge: a layer can lower its loss by inflating the scale of its weights rather than by separating the two populations. The analysis prescribes the repair - a whitened, scale-invariant goodness trained online within each layer - which we evaluate as a training procedure. Across three corpora, three depths and a fourfold range of layer width (13 seeds per cell), it raises linear-probe accuracy over the standard pairwise objective in every measured cell - by 4 to 7 points on eight of nine corpus-depth combinations - and closes 16-61% of the gap to end-to-end backpropagation. A control isolates the mechanism: Hinton's fixed-threshold loss also bounds the runaway, to a factor of 1.4 against 133, yet tracks the unmodified baseline - invariance to the gauge, not a bound on it, is what pays. Against the strongest published alternative - a sparse, top-k goodness - the derived objective is statistically indistinguishable on two corpora of three, yet only it removes the runaway: sparsity and gauge-invariance are independent axes, and the published variant recovers accuracy while leaving the pathology in place. We state the boundaries we measured, and every prediction was recorded before its experiment with the refutations reported.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑