arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经细胞自动机在其隐藏通道中学习通用特征

Neural Cellular Automata Learn General Features in their Hidden Channels

Etienne Guichard, Stefano Nichele

arXiv 2609.21870首次发表:更新:

发表机构

Østfold University of Applied Sciences(东福尔郡应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究神经细胞自动机隐藏通道的内部动态,提出将预训练教师隐藏状态注入学生模型的迁移学习机制,在少样本MNIST上以约9800参数超越循环和前馈架构,证明隐藏通道捕获通用尺度不变拓扑特征,实现参数高效迁移学习。

AI 中文摘要

现代深度学习模型通过过度参数化实现了令人印象深刻的泛化能力,但这一范式在少样本场景中常常面临过拟合和记忆化的问题。神经细胞自动机(NCAs)提供了一种高度参数高效的替代方案,然而现有研究主要关注其输出,对其内部隐藏通道的作用在很大程度上尚未探索。在本文中,我们研究了NCA隐藏通道的内部动态,并引入了一种新颖的迁移学习机制,该机制将预训练教师模型的隐藏状态注入学生模型,以引导早期优化。在少样本和尺度变化的MNIST基准测试上,NCAs优于可比较的循环和前馈架构,以极小的参数预算(约9,800个参数)展示了卓越的泛化能力。机制分析揭示,隐藏通道通过吸收形态复杂性和收敛到相互正交的状态,将特征提取与统一分类共识解耦。此外,我们证明这些隐藏通道捕获的是通用的、尺度不变拓扑基元,而非类别特定的模板。这使得学生模型能够利用仅在一部分数字(0-5)上训练的教师模型所迁移的特征,在未见类别上实现强大的少样本性能。我们的结果凸显了利用隐藏状态动态作为参数高效迁移学习的稳健、去中心化计算基质的潜力。

英文摘要

Modern deep learning models achieve impressive generalization through over-parameterization, but this paradigm often struggles with overfitting and memorization in few-shot regimes. Neural Cellular Automata (NCAs) offer a highly parameter-efficient alternative, yet research has focused primarily on their output, leaving the role of their internal hidden channels largely unexplored. In this paper, we investigate the internal dynamics of NCA hidden channels and introduce a novel transfer-learning mechanism that injects a pretrained teacher's hidden states into a student model to guide early optimization. Evaluated on few-shot and scale-variant MNIST benchmarks, NCAs outperform comparable recurrent and feed-forward architectures, demonstrating superior generalization with a minimal parameter budget (~9,800 parameters). Mechanistic analysis reveals that the hidden channels decouple feature extraction from uniform classification consensus by absorbing morphological complexity and converging to mutually orthogonal states. Furthermore, we demonstrate that these hidden channels capture general, scale-invariant topological primitives rather than class-specific templates. This allows a student model to achieve strong few-shot performance on unseen classes using features transferred from a teacher trained only on a subset of digits (0-5). Our results highlight the potential of utilizing hidden-state dynamics as a robust, decentralized computational substrate for parameter-efficient transfer learning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑