arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于群对比前向算法的涌现分层单语义神经元

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm

Yiming Tang, Qinglin Qi, Zhaoqian Yao, Harshvardhan Saini, Dianbo Liu

arXiv 2607.16295首次发表:更新:

发表机构

National University of Singapore; Lund University; Chinese University of Hong Kong; Indian Institute of Technology, Dhanbad(新加坡国立大学; 隆德大学; 香港中文大学; 印度理工学院(丹巴德分校))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨神经网络表示的可解释性,针对稀疏字典学习范式的局限,提出群对比前向算法GCFF,通过架构约束实现单语义性,能捕捉非线性概念,在CLIP表示上表现良好,还能从头训练网络并在图像分类基准中达最优性能。

AI 中文摘要

机械可解释性在理解神经网络表示方面取得了显著进展,稀疏字典学习(SDL)方法是核心范式,但存在局限性。我们假设存在不同的单语义性途径,生物视觉系统有高度选择性的神经元分层组织,源于局部、逐层学习规则。为此提出群对比前向算法(GCFF),通过架构约束而非稀疏性实现单语义性,能捕捉非线性概念。在CLIP表示上,单个训练的GCFF模块可恢复抽象度随深度递增的单语义神经元,且无需稀疏约束或抽象级别监督。此外,GCFF能从头训练网络,在各种图像分类基准上达到前向算法的最优性能。

英文摘要

Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse autoencoders, as a central paradigm. However, recent work has reported several limitations of this paradigm: SDL objectives are non-identifiable; SDL methods rely heavily on the Linear Representation Hypothesis; and a growing body of evidence points to concepts that are encoded non-linearly and are therefore not expressible as any single direction. We hypothesise that a different route to monosemanticity is available. Biological visual systems exhibit highly selective neurons organised into hierarchies of increasing abstraction, and this organisation emerges from local, layer-wise learning rules rather than from a global error signal; we therefore ask whether a biologically plausible learning algorithm will likewise yield monosemantic neurons. To test this, we propose Group-Contrastive Forward-Forward (GCFF), a forward-forward training algorithm that combines class-specific routing with within-class contrastive objectives, reaching monosemanticity through architectural constraints rather than sparsity. Because GCFF attaches multiple non-linear layers to the representation under study, its neurons can therefore capture the non-linear concepts. On CLIP representations, a single trained GCFF module recovers monosemantic neurons whose abstraction increases progressively with depth, reaching environmental properties that hold independently of an image's foreground, without any sparsity constraint or supervision of abstraction level. We further demonstrate that GCFF can train networks from scratch, achieving state-of-the-art performance among forward-forward algorithms on various image classification benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑