任务相关零空间残差用于非单射神经映射
Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings
浏览论文内容
中文总结 AI 辅助
提出任务相关零空间残差(NSR)框架,通过提取非单射线性映射的零空间分量并集成到下游任务,在令牌合并和图聚合中显著提升性能,弥补映射导致的输入信息丢失。
中文摘要 AI 辅助
神经网络中的非单射映射将不同的输入映射到相同的表示,从而在输入空间中隐式地诱导出等价关系。然而,这些映射所消除的输入差异可能仍然是下游任务所需要的,这造成了算子诱导的不可区分性与任务所需的区分性之间的不匹配。对于在当前前向传播中实现的非单射线性算子,其零空间恰好刻画了这些不可见的输入变化。我们提出了任务相关零空间残差(NSR),这是一个用于非单射线性映射的通用残差框架。NSR结合了从映射前表示中提取零空间分量、成员级编码与门控,以及特定应用集成,在下游监督下利用潜在的任务相关信息,同时保留原始的聚合或合并规则。我们在两种结构不同的设置中评估了NSR:令牌合并和图聚合。在令牌合并中,NSR在36个评估配置中的34个中取得了比相应压缩基线更高的语义分割性能,在强压缩下最大观测增益为31.51 mIoU点。在图聚合中,NSR在三个骨干网络上,在深度d=2--6的Tree-NeighborsMatch上实现了100%的训练准确率,同时在异质性节点分类和分子图回归上也取得了增益。这些结果共同支持零空间残差作为非单射线性映射的实用补充,使下游模型能够从原始算子输出中不可见的输入区分中学习。
英文摘要
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-induced indistinguishability and task-required distinctions. For non-injective linear operators realized in the current forward pass, their null spaces exactly characterize these invisible input variations. We propose Task-Relevant Null-Space Residuals (NSR), a general residual framework for non-injective linear mappings. NSR combines null-space component extraction from pre-mapping representations, member-level encoding and gating, and application-specific integration to exploit potentially task-relevant information under downstream supervision while preserving the original aggregation or merging rules. We evaluate NSR in two structurally different settings: token merging and graph aggregation. In token merging, NSR achieves higher semantic segmentation performance than the corresponding compressed baselines in 34 out of 36 evaluated configurations, with a maximum observed gain of 31.51 mIoU points under strong compression. In graph aggregation, NSR achieves 100% training accuracy on Tree-NeighborsMatch at depths d=2--6 across three backbones, alongside gains on heterophilic node classification and molecular graph regression. Together, these results support null-space residuals as a practical complement to non-injective linear mappings, enabling downstream models to learn from input distinctions invisible in the original operator's output.
发表机构
- Institute of Artificial Intelligence Innovation and Industry, Fudan University(复旦大学人工智能创新与产业研究院)
- Shanghai Academy of AI for Science(上海人工智能科学研究院)
- Human Phenome Institute, Fudan University(复旦大学人类表型组研究院)
- School of Information and Communication Engineering, Communication University of China(中国传媒大学信息与通信工程学院)
机构由 AI 辅助整理,请以论文原文为准。