arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40258cs.NE

大语言模型引导的脉冲序列建模原生神经架构进化发现

Large Language Model-Guided Evolutionary Discovery of Native Neural Architectures for Spiking Sequence Modeling

Ruoyu Zhao, Jiaqi Wu, Chenyu Zhu, Zhichao Lu

首次发表
浏览论文内容

中文总结 AI 辅助

提出OpenArchEvo,利用大语言模型在开放程序空间中进化原生脉冲神经网络架构,通过三视图表示预测性能与新颖性,发现NeuroGate等架构,超越ANN基线并大幅降低能耗。

中文摘要 AI 辅助

脉冲神经网络(SNN)通过稀疏、事件驱动的计算实现低能耗序列建模。然而,脉冲编码、神经元动力学和信息传播之间的相互作用使架构设计变得复杂。现有的SNN序列模型通常采用为实值激活设计的人工神经网络(ANN)架构,可能未能充分利用基于脉冲的通信和时态状态更新,这促使了原生SNN架构的自动发现。大多数进化神经架构搜索(ENAS)方法在预定义的配置空间内运行,将发现限制在那些空间内可表达的机制中。我们提出了OpenArchEvo,它利用大语言模型(LLM)在脉冲投影约束下的开放程序空间中进化可执行的架构代码。在该空间中,代码差异不一定反映架构新颖性,而直接性能评估需要昂贵的训练。我们构建了一个涵盖代码、设计原理和行为指纹的三视图表示,以支持新颖性估计和性能预测。搜索将预测性能和估计新颖性作为两个目标,使用代理预测来选择候选进行昂贵的训练评估。估计候选训练成本为132个V100 GPU天,搜索发现了多种原生SNN架构,以三种设计为例,其机制包括脉冲活动依赖的状态更新控制和残差路径。所发现的NeuroGate在WikiText-103上超越了ANN DeltaNet,并且所发现的架构相对于常见的密集Transformer(ANN)基线,将估计的架构级算术能耗降低了最多50.6倍(LoopMem)。所有代码和所有发现的架构将很快公开提供。

英文摘要

Spiking neural networks (SNNs) offer low-energy sequence modeling through sparse, event-driven computation. However, interactions among spike encoding, neuronal dynamics, and information propagation complicate architecture design. Existing SNN sequence models often adapt artificial neural network (ANN) architectures designed for real-valued activations, potentially underusing spike-based communication and temporal state updates, motivating automated discovery of native SNN architectures. Most evolutionary neural architecture search (ENAS) methods operate within predefined configuration spaces, limiting discovery to mechanisms expressible within those spaces. We introduce OpenArchEvo, which uses large language models (LLMs) to evolve executable architecture code in an open program space under spiking-projection constraints. In this space, code differences need not reflect architectural novelty, while direct performance evaluation requires costly training. We construct a three-view representation spanning code, design rationale, and a behavioral fingerprint to support novelty estimation and performance prediction. The search treats predicted performance and estimated novelty as two objectives, using surrogate predictions to select candidates for expensive training evaluations. With an estimated candidate-training cost of 132 V100 GPU-days, the search uncovers multiple native SNN architectures, exemplified by three designs featuring mechanisms such as spike-activity-dependent control of state updates and residual pathways. The discovered NeuroGate surpasses the ANN DeltaNet on WikiText-103, and the discovered architectures reduce estimated architecture-level arithmetic energy by up to 50.6x (LoopMem) relative to a common dense Transformer (ANN) baseline. All code and all discovered architectures will be made publicly available soon.

发表机构

  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑