arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经元汤:无需反向传播进化异步共享神经元时间图

NeuronSoup: Evolving Asynchronous, Shared-Neuron Temporal Graphs without Backpropagation

Subodh Kalia

arXiv 2607.15217首次发表:更新:

发表机构

Subodh Kalia(Subodh Kalia)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出神经元汤架构,以异步共享神经元信号传播取代同步逐层处理。通过遗传算法共同进化架构各要素,在MNIST数字分类任务中,利用冻结的ResNet18特征输入,进化出特定网络,达到较高准确率,解决了当前深度学习的一些局限。

AI 中文摘要

我们提出了神经元汤,这是一种神经计算架构,它通过共享神经元池进行异步、延迟介导的信号传播来取代同步的逐层处理。网络中的每条路径通过可变数量的中间隐藏神经元将连续值信号从一个输入神经元路由到一个输出神经元。隐藏神经元在路径间物理共享,当两条路径通过同一神经元时,第二个到达的信号会遇到第一个留下的累积状态,产生取决于信号极性和到达时间的相长或相消干涉。整个架构——拓扑、权重、延迟和连接性——由对14602个基因的实值基因组进行操作的遗传算法共同进化。在使用冻结的ResNet18特征作为输入的10类MNIST数字分类任务中,该系统通过266个隐藏神经元(其中156个在多条路径间共享,一个神经元参与11条不同路径)进化出一个包含204条活跃路径的网络,在10000代后达到85.9%的测试准确率,训练后的模型占用115KB。我们认为该架构解决了当前深度学习的基本局限性:它不需要可微计算图,能针对每个样本调整计算深度,并发现当前架构必须明确设计的处理路径间的横向相互作用。我们还讨论了为什么遗传算法是解决此类问题的正确优化工具,为什么CMA-ES在这个规模上失败,以及该架构如何通过替换编码器和输出结构推广到任意领域。

英文摘要

We present NeuronSoup, a neural computation architecture that replaces synchronous layer-by-layer processing with asynchronous, delay-mediated signal propagation through a pool of shared neurons. Each path in the network routes a continuous-valued signal from one input neuron to one output neuron through a variable number of intermediate hidden neurons. Hidden neurons are physically shared across paths: when two paths pass through the same neuron, the second arrival encounters the accumulated state left by the first, producing constructive or destructive interference that depends on signal polarity and arrival timing. The entire architecture -- topology, weights, delays, and connectivity -- is co-evolved by a genetic algorithm operating on a flat real-valued genome of 14,602 genes. On 10-class MNIST digit classification using frozen ResNet18 features as input, the system evolves a network of 204 active paths through 266 hidden neurons (156 shared across multiple paths, with one neuron participating in 11 distinct paths) and achieves 85.9\% test accuracy after 10,000 generations. The trained model occupies 115 KB. We argue that this architecture addresses fundamental limitations of current deep learning: it requires no differentiable computation graph, adapts its computation depth per-sample, and discovers lateral interactions between processing pathways that current architectures must engineer explicitly. We discuss why genetic algorithms are the correct optimization tool for this problem class, why CMA-ES fails at this scale, and how the architecture generalizes to arbitrary domains by substituting the encoder and output structure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑