局部性的代价:为何前向-前向算法不如反向传播?
The Price of Locality: Why Forward-Forward Underperforms Backpropagation?
AI总结:
本文诊断前向-前向算法相对反向传播的性能差距,指出其因局部更新导致优化下限和表示坍缩,从而限制了与深度无关的学习容量。
AI中文摘要:
前向-前向算法(FFA)用逐层的局部对比目标取代反向传播(BP),从而消除了反向传递以及对中间激活值进行保留的需求,但其与BP之间持续存在的性能差距会随着深度增加而恶化。本文诊断出两种结构性失效模式:一是由并发局部更新引发的优化下限;二是由局部更新机制导致的层表示几何坍缩。在优化方面,我们证明了FFA的损失在每一层都满足Polyak--Lojasiewicz不等式;然而,各层同时更新会引发层间表示分布漂移,因此每一层都在针对一个移动的输入分布进行优化,并产生误差下限。在表示方面,层表示的两两相似性核随深度增加以指数方式收缩至秩一,导致逐层误差信号的多样性坍缩。这种坍缩使得FFA的有效学习容量(衡量各层梯度信息多样性的指标)与深度无关,而BP的链式法则信号则保留了逐层的多样性,从而获得随深度扩展的容量。
英文摘要:
The Forward-Forward Algorithm (FFA) replaces backpropagation (BP) with layer-wise local contrastive objectives, eliminating the backward pass and the need to retain intermediate activations, yet suffers a persistent performance gap with BP that worsens with depth. This paper diagnoses two structural failure modes: an optimization floor arising from concurrent local updates; and a geometric collapse of layer representations driven by the local update mechanism. On the optimization side, we prove that the FFA loss satisfies the Polyak--Lojasiewicz inequality at each layer; however, simultaneous layer updates induce inter-layer representation-distribution drift, so each layer optimizes against a moving input distribution and incurs an error floor. On the representational side, the pairwise similarity kernel of layer representations contracts exponentially toward rank one as depth increases, collapsing the diversity of per-layer error signals. This collapse bounds FFA's effective learning capacity, which measures the diversity of gradient information across layers, independently of depth, whereas BP's chain-rule signal preserves per-layer diversity, yielding a capacity that scales with depth.