arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31589cs.LGcs.NE

直接反馈对齐中的共模坍缩与恢复

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Varun Reddy, Bernardo L. Sabatini, Houman Safaai

首次发表
浏览论文内容

中文总结 AI 辅助

本文揭示直接反馈对齐中误差共模导致tanh单元饱和坍缩,通过均值-协方差分解定位机制,提出校准读出和减去批次均值可抑制坍缩并加速学习。

中文摘要 AI 辅助

直接反馈对齐(DFA)通过输出误差的固定随机投影来训练隐藏层。对于tanh隐藏单元和独立的sigmoid输出,普通的随机梯度下降可能会在类别频率的常数预测器的损失附近停滞。我们将这种停滞归因于误差的共模,即输入间共享的分量。精确的均值-协方差分解将平均教学信号和平均突触前活动形成的秩一更新分离出来。其主要分量驱动tanh单元趋于饱和。在初始化时,随机反馈在平均意义上无法对共享误差进行系统性修正;读出学习限制了其持续时间。从网络初始化的简化模型(无拟合参数)在48种设置下预测了激活敏感性的集中。在MNIST上,类别可解码性在坍缩后基本保持,但在固定学习率下读出学习仍然缓慢。尽管坍缩更深,Adam学习更快。将基线读出校准到类别先验可抑制坍缩并加速学习;较弱的反馈以较慢的学习换取较少的坍缩。用误差的符号替换误差会维持坍缩;减去信号的批次均值可防止持续坍缩并改善测试设置中的学习。相关效应也出现在更深层和卷积网络以及CIFAR-10上,其严重性和代价取决于读出、优化器和输入统计。

英文摘要

Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learning remains slow at a fixed learning rate. Adam learns faster despite deeper collapse. Calibrating the baseline readout to the class prior suppresses collapse and speeds learning; weaker feedback trades less collapse for slower learning. Replacing errors by their signs sustains collapse; subtracting the signal's batch mean prevents sustained collapse and improves learning in the tested setting. Related effects occur in deeper and convolutional networks and on CIFAR-10, with severity and cost depending on the readout, optimizer and input statistics.

发表机构

  • Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University(哈佛大学肯普纳自然与人工智能研究所)
  • Howard Hughes Medical Institute(霍华德·休斯医学研究所)
  • Harvard Medical School(哈佛医学院)

机构由 AI 辅助整理,请以论文原文为准。

↑