发表机构
Fudan University; University of Washington; Tongji University; Shandong University; Datacanvas(复旦大学; 华盛顿大学; 同济大学; 山东大学; Datacanvas)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过莫比乌斯反演分解残差网络输出,发现其交互高阶主导且随输入变化,证明密集残差网络隐式实现软路由,无需显式路由器或稀疏执行。
AI 中文摘要
残差网络对每个输入都执行每个块,但其功能贡献未必与输入无关。我们将训练好的残差网络表述为关于二元残差分支掩码的集合函数,并应用莫比乌斯反演将其输出精确分解为单个残差修正项和高阶交互项。对于平滑残差堆栈,我们证明在残差缩放下,每个固定的 $k$ 路交互项按 $\mathcal{O}(\lambda^k)$ 缩放。对 ImageNet 预训练的 ResNet-18 和 ResNet-34 的详尽分析揭示,交互质量分别在五阶和十阶达到峰值,而非在低阶。减小残差缩放会将两个谱向低阶移动,但也会改变模型预测。交互系数在幅度上集中,但并非硬稀疏,且保持预测的稀疏性随深度减弱。关键在于,主导交互在不同输入和预测类别之间围绕一个共享的全局核心变化,而其整体复杂度随样本难度变化不大。这些结果表明,密集残差网络实现了一种隐式的软路由形式:每个块都被执行,但不同输入依赖于不同的残差交互专家。因此,路由可以在功能贡献层面出现,而无需显式路由器或稀疏执行。
英文摘要
Residual networks execute every block for every input, yet their functional contributions need not be input independent. We formulate a trained ResNet as a set function over binary residual-branch masks and apply Möbius inversion to decompose its output exactly into individual residual corrections and higher-order interactions. For smooth residual stacks, we show that each fixed $k$-way interaction scales as $\mathcal{O}(λ^k)$ under residual scaling. Exhaustive analysis of ImageNet-pretrained ResNet-18 and ResNet-34 reveals that interaction mass peaks at orders five and ten, respectively, rather than at low orders. Reducing the residual scale shifts both spectra toward lower orders, but also changes model predictions. The interaction coefficients are concentrated in magnitude but not hard sparse, and prediction-preserving sparsity weakens with depth. Crucially, the dominant interactions vary across inputs and predicted classes around a shared global core, while their overall complexity changes little with sample difficulty. These results show that dense ResNets implement an implicit form of soft routing: every block is executed, but different inputs rely on different residual interaction experts. Routing can therefore emerge at the level of functional contribution without an explicit router or sparse execution.
Comments46 pages, 14 figures, 18 tables