发表机构
Microsoft; Technion – Israel Institute of Technology; NVIDIA(微软; 以色列理工学院; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究有限探针表示对神经泛函学习的充分性,建立辨识性与普适性理论,提出HIDDENPROBE架构,在MLP和Transformer基准上达到最先进性能。
AI 中文摘要
近年来,学习神经网络的属性引起了越来越多的关注,现有方法要么直接作用于网络参数,要么通过基于探针的网络行为表示。虽然探针方法表现出强大的经验性能,但其理论基础仍然有限。在这项工作中,我们研究了有限基于探针的表示何时足以用于学习神经泛函。我们建立了探针的一般辨识性和普适性结果,并表明使用中间隐藏表示可以比仅依赖最终输出提供更具信息量的表示。受这些结果的启发,我们引入了HIDDENPROBE,一种用于从隐藏探针响应中学习的简单架构。在一系列神经泛函基准测试中,包括MLP和Transformer,HIDDENPROBE持续优于现有探针方法,并实现了最先进的性能。我们的代码已在GitHub上公开。
英文摘要
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical foundations remain limited. In this work, we study when finite probe-based representations are sufficient for learning neural functionals. We establish general identification and universality results for probing, and show that using intermediate hidden representations can provide significantly more informative representations than relying only on final outputs. Motivated by these results, we introduce HIDDENPROBE, a simple architecture for learning from hidden probe responses. Across a range of neural functional benchmarks, including both MLPs and Transformers, HIDDENPROBE consistently improves over existing probing methods and achieves state-of-the-art performance. Our code is publicly available on GitHub.