发表机构
Institut quantique, Université de Sherbrooke(谢布鲁克大学量子研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出知识矩阵作为前馈网络输入的更高表示,证明其能恢复实现芽并归因于梯度,且提供跨架构的对抗性度量,但存在局限性。
AI 中文摘要
我们研究训练好的前馈网络的知识矩阵,将其作为其输入的更高表示。网络是一对 $(W,f)$,即其箭图的薄表示 $W$ 和激活函数 $f$;其函数通过箭图表示的空间进行分解,每个输入 $x$ 诱导一个表示,知识矩阵 $M(x)\in\mathbb{R}^{C\times(d+1)}$ 是该表示收缩为一个矩阵,其行和恰好等于 logits。对于一个训练好的网络,我们问什么决定了它,它对什么不变,它决定了什么,以及它的几何度量了什么。在 (LCS)(局部常数斜率对角线,如 ReLU)下,规则输入处的矩阵是实现芽的函数;其在规则编码中的稳定子恰好是输入无消失坐标时的芽稳定子,神经元置换是特例;并且它能恢复芽,而隐藏激活(规范协变且芽不完全)则不足够。在 (LCS) 下,它等于每类梯度乘以输入加上精确的聚合偏置归因,将其扎根于归因理论,并通过 $C$ 个向量-雅可比积计算,而非探测。固定形状给出了 ResNet-152、DenseNet-121 和 GoogLeNet 之间无需对齐的每样本距离;行和恒等式给出了精确的可见/不可见位移分解,其无单位相干性 $A=(d_\Psi/d_M)^2$ 将对抗性芽运动置于中位数 $A\le 0.23$,攻击族排序在六种架构上一致(Kendall $W=0.921$;在三种全尺寸网络上为 $0.97$)。两个诚实的负面结果:在 AlexNet/CIFAR-10 上,倒数第二层特征赢得 6 个检测器中的 5 个和全部 16 个攻击,而矩阵方向反事实在 0/54 上失败。
英文摘要
We study the knowledge matrix of a trained feedforward network as a higher representation of its inputs. A network is a pair $(W,f)$, a thin representation $W$ of its quiver and an activation $f$; its function factorizes through the space of quiver representations, each input $x$ inducing a representation, and the knowledge matrix $M(x)\in\mathbb{R}^{C\times(d+1)}$ is the contraction of that representation to one matrix whose rows sum exactly to the logits. At one trained network we ask what determines it, what it is invariant to, what it determines, and what its geometry measures. Under (LCS), a locally constant slope diagonal, as for ReLU, the matrix at a regular input is a function of the realized germ; its stabilizer among encodings regular there is exactly the germ stabilizer at inputs with no vanishing coordinate, neuron permutation a special case; and it recovers the germ, whereas hidden activations, gauge-covariant and germ-incomplete, are not enough. Under (LCS) it equals per-class gradient$\times$input plus an exact aggregate bias attribution, grounding it in attribution theory and computing it by $C$ vector-Jacobian products instead of probing. The fixed shape gives an alignment-free per-sample distance between ResNet-152, DenseNet-121 and GoogLeNet; the row-sum identity gives an exact visible/invisible displacement decomposition whose unit-free coherence $A=(d_Ψ/d_M)^2$ puts adversarial germ motion at median $A\le 0.23$, with an attack-family ordering concordant across six architectures (Kendall $W=0.921$; $0.97$ on the three networks at full scale). Two honest negatives: on AlexNet/CIFAR-10 penultimate features win 5 of 6 detectors and all 16 attacks, and a matrix-direction counterfactual fails 0/54.
Comments86 pages, main text 35 pages, appendices 51 pages