arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

增加宽度使得贪心逐层训练在自监督学习中能与端到端反向传播相媲美

Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning

Syon Mansur, Joel Zylberberg

arXiv 2610.00753首次发表:更新:

发表机构

Jules Stein Eye Institute; University of California, Los Angeles(朱尔斯·斯坦眼科研究所; 加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨自监督学习中网络宽度与深度对贪心逐层训练和端到端反向传播的影响,发现宽网络中贪心训练性能可媲美甚至超越端到端训练,并揭示表征几何差异为潜在机制。

AI 中文摘要

端到端反向传播一直是深度学习中的主导训练模式,它允许神经网络各层之间的参数更新进行协调。先前的研究探索了替代的——在某些情况下更简单的——训练机制,表明它们有时能达到与反向传播相似的性能。然而,在何种架构条件下,避免误差端到端反向传播的局部优化网络能够学习到与端到端训练所学的表征相媲美的表征,仍不清楚。我们旨在自监督学习的背景下回答这个问题,自监督学习是人工智能中大规模预训练的重要框架。在此,我们研究了网络宽度和深度如何影响卷积网络中贪心逐层训练和端到端自监督训练的有效性。我们发现,在更宽的网络中,端到端反向传播相对于贪心逐层训练的优势会缩小:在相对较浅且非常宽的网络中,我们甚至观察到使用贪心逐层训练训练的模型性能更高。随后对这些网络形成的表征的分析表明,非常宽的贪心训练网络比用反向传播端到端训练的网络表现出更有利的表征几何。这项工作表明,宽度可以弥补受限的信用分配,并识别出表征几何的差异作为其性能提升的潜在机制。

英文摘要

End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑