arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06862cs.LG

神经网络中的特征叠加:从理论到实践

Feature Superposition in Neural Networks: From Theory to Practice

Dai Shi, Xiaoyu Li, Andi Han, José Miguel Hernández-Lobato

首次发表
浏览论文内容

中文总结 AI 辅助

本文综述了神经网络中特征叠加现象的理论与实证研究,比较了特征恢复方法,并指出了开放问题,旨在推动对叠加的更深入理解和更可靠的网络解释方法。

中文摘要 AI 辅助

叠加是指神经网络表示的特征数量超过其维度数量的现象。它为多义神经元提供了可能的解释,并推动了从神经激活中恢复可解释特征的方法。理论模型通常从一组给定的输入特征以及这些特征值在不同输入间如何变化的假设出发,然后研究网络如何在低维隐藏表示中编码这些值。相比之下,实证工作旨在识别训练网络中编码的特征,并确定它们在计算中的作用。在本综述中,我们回顾了叠加表示的几何、学习和计算,解释了特征统计和解码器选择如何影响结论。为了将这些理论解释与训练网络中的证据联系起来,我们比较了恢复和分析特征的实用方法,并考察了它们的评估所确立的内容。由于仅凭准确的激活重建并不能确立特征身份或因果用途,我们根据这些不同主张可获得的证据,讨论了这些方法已记录在案的失败和应用。最后,我们评估了先前提出的开放问题,并确定了关于训练网络中叠加的剩余理论和实证问题。我们希望我们的工作能为更深入地理解叠加以及更可靠的神经网络解释方法铺平道路。

英文摘要

Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptions about how their values vary across inputs, then study how a network encodes those values in a lower-dimensional hidden representation. Empirical work, by contrast, seeks to identify the features encoded in trained networks and determine their role in computation. In this survey, we review the geometry, learning, and computation of superposed representations, explaining how feature statistics and decoder choice affect the conclusions. To connect these theoretical accounts with evidence from trained networks, we compare practical methods for recovering and analyzing features and examine what their evaluations establish. Since accurate activation reconstruction alone does not establish feature identity or causal use, we discuss the methods' documented failures and applications in light of the evidence available for these different claims. Finally, we assess previously stated open problems and identify remaining theoretical and empirical questions about superposition in trained networks. We hope our work can pave the way for a deeper understanding of superposition and more reliable methods for interpreting neural networks.

发表机构

  • University of Cambridge(剑桥大学)
  • University of New South Wales(新南威尔士大学)
  • University of Sydney(悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

↑