arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22554cs.LGcs.AI

反向传播权重的起起落落

The Ups and Downs of Backprop Weights

Giuseppe Chindemi, Benjamin F. Grewe

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出权重算子概念,通过两阶段学习实现功能组件的可重用与选择性更新,并以向量网络验证,旨在解决权重纠缠问题,促进功能参数可识别性。

中文摘要 AI 辅助

反向传播(BP)通过使大型分层网络能够端到端地学习复杂函数,推动了现代深度学习的显著成功。然而,它本身并未决定应如何组织参数,以便功能组件能够被重用并有选择性地调整。例如,物体识别和运动预测可能依赖于重叠的参数集,这使得它们难以被独立隔离或修改。我们将这种情况称为权重纠缠。现代架构动态地选择网络的哪些部分来处理每个样本:非线性门控单元、注意力选择交互,以及专家混合架构将输入路由到模块。然而,这种选择并不能确保相同的功能组件在样本间始终与可识别的参数集相关联。我们提出了权重算子:参数化模块,实现可重用的功能组件,并可在推理时组合以形成每个样本所需的函数。学习分两个阶段进行:模型首先推断所需的算子组合,然后仅更新所选算子的参数集。向量网络(VNs)提供了一种实现。它们将算子选择与每层内的局部误差驱动更新相结合,并表明学习到的算子可以在训练中未出现的组合中被重用,同时更新仍仅限于所选参数集。这为测试功能参数可识别性提供了基础:即在学习过程中,算子是否始终与相同的功能组件相关联。我们认为,功能参数可识别性可能为模型提供一种组织原则,使模型能够系统地重用和重组学习到的函数,同时仅调整需要改变的组件。

英文摘要

Backpropagation (BP) has driven the remarkable success of modern deep learning by enabling large hierarchical networks to learn complex functions end-to-end. Yet it does not by itself determine how parameters should be organized so that functional components can be reused and adapted selectively. For example, object recognition and motion prediction may depend on overlapping parameter sets, making them difficult to isolate or modify independently. We call this condition weight entanglement. Modern architectures dynamically select which parts of a network process each sample: nonlinearities gate units, attention selects interactions, and Mixture-of-Experts architectures route inputs to modules. Yet such selection does not ensure that the same functional component remains linked to an identifiable parameter set across samples. We propose weight operators: parameterized modules that implement reusable functional components and can be composed at inference to form the function required by each sample. Learning proceeds in two stages: the model first infers the required operator composition, then updates only the selected operators' parameter sets. Vector Networks (VNs) provide one implementation. They couple operator selection to local error-driven updates within each layer and show that learned operators can be reused in combinations absent from training while updates remain restricted to the selected parameter sets. This provides a basis for testing functional parameter identifiability: whether an operator remains linked to the same functional component during learning. We argue that functional parameter identifiability may provide an organizing principle for models that systematically reuse and recombine learned functions while adapting only the components that need to change.

发表机构

  • Institute of Neuroinformatics, UZH / ETH Zurich(苏黎世大学/苏黎世联邦理工学院神经信息学研究所)
  • ETH AI Center(苏黎世联邦理工学院人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

↑