发表机构
University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现卷积序列标注器的归一化层可通过序列跨越路径提供全局上下文,在标签长序列场景中可替代大部分感受野的作用,且消融感受野模块会高估其贡献。
AI 中文摘要
卷积序列标注器的感受野通常被视为模型可用上下文的范围:它设定膨胀率调度、限制流式处理的视界,并为局部性主张提供依据。然而,我们证明这可能是错误的:当归一化层在推理时沿序列从当前输入计算统计量时,这些统计量会打开一条跨越序列的路径,绕过卷积感受野以提供全局上下文。我们从该层的雅可比矩阵推导出这一结论(该判据无需实验),且该路径所承载的内容具有闭式形式。在具有可计算最优值的合成标注过程中,当标签连续出现长序列时,序列跨越归一化编码的全局摘要已能提供更大感受野所能带来的几乎全部收益:达到9个位置的网络与全序列最优值的差距仅为0.009,而其感受野范围对应的准确率接近随机水平。通过对每个位置计算相同统计量来关闭该路径,在所有测试难度水平的模拟基因组以及真实的1000 Genomes单倍型数据上,将增大感受野的价值提升了最高一个数量级。该路径还会混淆归因:对已训练网络中增大感受野的模块进行消融会切断部分路径,相对于从头重新训练,会高估这些模块的贡献8.3至16.1倍。当标签切换更频繁时,归一化对感受野的替代作用会减弱。在标签连续出现长序列的场景中,感受野的依据和消融分析的数值都没有错误,但两者都将功劳归于了错误的组件。
英文摘要
A convolutional sequence labeler's receptive field is routinely treated as the extent of the model's usable context: it sets dilation schedules, bounds streaming horizons, and underwrites locality claims. However, we show that this can be false: when a normalization layer computes statistics from the current input along the sequence at inference, those statistics open a sequence-spanning path that bypasses the convolutional receptive field to provide global context. We derive this from the layer's Jacobian (the criterion needs no experiment), and what the path carries has a closed form. On a synthetic labeling process with computable optima, the global summary that a sequence-spanning normalization encodes already supplies almost all of what a larger receptive field would buy where labels come in long runs: a network reaching 9 positions comes within 0.009 of the whole-sequence optimum, against a near-chance bound for its reach. Closing the path, by taking the same statistics per position, multiplies what enlarging the receptive field is worth by up to an order of magnitude on simulated genomes at every difficulty level tested and on real 1000 Genomes haplotypes. The same path also confounds attribution: ablating a trained network's receptive-field-enlarging blocks severs part of the path, overstating their contribution 8.3-16.1-fold relative to retraining from scratch. The substitution of normalization for receptive field fades as labels switch more often. Where labels run long, neither the receptive-field justification nor the ablation is wrong about its numbers, but both credit the wrong component.