AI 中文总结
研究如何解释神经网络内部特征从特定输入中提取的信息,提出源接地特征反演方法,通过特定条件和映射修复信号,经DAG反向传递组合状态,支持多种架构,验证反演依赖关系,打开模型隐藏特征层次结构。
AI 中文摘要
解释神经网络需要理解其内部特征从特定输入中提取了什么。特征反演旨在在输入域中表达选定特征,但传统迭代方法搜索的是其重新编码表示与目标匹配的输入。由于许多输入可满足此约束,仅目标匹配无法指定与生成该特征的样本相关的反演。我们通过以目标生成输入处的源局部网络几何为条件来制定源接地特征反演。在计算DAG的每个边界,反向传播提供正确的反向依赖,但传输伴随信号而非上游状态估计。我们用从均值种子VJP到上游状态的闭式矩阵维纳映射局部修复此信号,再用第二个维纳映射处理JVP前向一致性残差,并通过同一DAG在一次有限反向传递中组合修复后的状态。一个校准零截距映射族支持跨多种CNN和Transformer架构、张量组件及视觉分布的新输入、深度、通道和通道组,无需特定查询优化。匹配的目标和源控制验证每个反演依赖于选定特征和被解释样本的局部算子,而非与目标无关的图像模板。预测条件特征图集将这些可视化与对相应内部特征的独立干预对齐。源接地特征反演共同打开了模型隐藏特征层次结构,以便在单个层和通道级别进行检查,将网络从输入中提取的内容与塑造其决策的内部证据联系起来。
英文摘要
Interpreting a neural network requires understanding what its internal features extract from a particular input. Feature inversion seeks to express a selected feature in the input domain, but canonical iterative methods search for an input whose re-encoded representation matches the target. Because many inputs can satisfy this constraint, target matching alone does not specify the inverse associated with the sample that generated the feature. We formulate source-grounded feature inversion by conditioning the inverse on the source-local network geometry at the target-generating input. At each boundary of the computational DAG, backpropagation provides the correct reverse dependencies but transports an adjoint signal rather than an upstream-state estimate. We locally repair this signal with a closed-form matrix Wiener map from a mean-seed VJP to the upstream state, followed by a second Wiener map for the JVP forward-consistency residual, and compose the repaired states through the same DAG in one finite reverse pass. One calibrated zero-intercept map family supports new inputs, depths, channels, and channel groups across diverse CNN and Transformer architectures, tensor components, and visual distributions without query-specific optimisation. Matched target and source controls verify that each inverse depends on the selected feature and the local operators of the sample being explained, rather than a target-independent image template. Prediction-conditioned feature atlases align these visualisations with independent interventions on the corresponding internal features. Together, source-grounded feature inversion opens the model's hidden feature hierarchy to inspection at the level of individual layers and channels, linking what the network extracts from an input to the internal evidence that shapes its decision.