arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过分裂前向梯度进行无反向传播的主干训练

Backpropagation-Free Trunk Training via the Split Forward Gradients

Tian Qin, Wei-Min Huang

arXiv 2607.16612首次发表:更新:

发表机构

Lehigh University(莱维大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对反向传播训练深度网络内存密集问题,提出Split-FG方法,在中间表示处拆分网络,精确计算输出头梯度,估计主干梯度,降低估计器方差,无需反向传播主干,在多个基准测试中取得较好结果并减少内存。

AI 中文摘要

反向传播使深度网络训练内存密集,因为要存储中间激活值。前向模式方法避免此成本,但随着训练参数数量增加,梯度估计噪声增大。我们引入分裂前向梯度(Split-FG),在中间表示处拆分网络,精确计算输出头梯度,仅用雅可比向量积估计主干梯度。这降低估计器方差,无需反向传播主干,还保留了类似Adam的收敛保证。实验揭示了一个实际失败模式。在WikiText-103上,主干的朴素前向梯度训练比冻结随机初始化的主干更差,可能是Adam对每个噪声大、欠定的主干坐标更新过激进。简单为主干使用小得多的学习率可扭转此结果。Split-FG在表格基准测试中也产生了最强的无反向传播结果,在CIFAR-10上达到60.5%,在CIFAR-100上达到35.2%。相对于匹配的反向传播,它可将峰值内存减少多达35%,尽管随着前向模式主干增长,性能差距会扩大。

英文摘要

Backpropagation makes training deep networks memory intensive because it must store intermediate activations. Forward-mode methods avoid this cost, but their gradient estimates become increasingly noisy as the number of trained parameters grows. We introduce Split Forward Gradient (Split-FG), which splits a network at an intermediate representation: it computes the output head gradient exactly and estimates only the trunk gradient with a Jacobian--vector product. This reduces estimator variance and requires no backward pass through the trunk, while retaining an Adam-style convergence guarantee. Our experiments reveal an important practical failure mode. On WikiText-103, naive forward-gradient training of the trunk performs worse than leaving a randomly initialized trunk frozen, likely because Adam updates every noisy, under-determined trunk coordinate too aggressively. Simply using a much smaller learning rate for the trunk reverses this result: a $16$M-parameter GPT-2-style model reaches validation perplexity $387$, compared with $668$ for the frozen-trunk control and $2{,}885$ for a matched pure forward-gradient baseline (backpropagation reaches $150$). Split-FG also produces the strongest backprop-free results on our tabular benchmarks and reaches $60.5\%$ on CIFAR-10 and $35.2\%$ on CIFAR-100 with a heavy-head design. It reduces peak memory by up to $35\%$ relative to matched backpropagation, although the performance gap widens as the forward-mode trunk grows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑