arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25088cs.LGcs.AIq-bio.NC

用于神经解码的冯·诺依曼状态空间Transformer

The Von-Neumann State-Space Transformer for neural decoding

Morteza Sarafyazd

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出VN-SST模型,其前馈模块为低维指令库,在神经解码任务中样本与参数效率优于现有模型,且在语言建模任务中也展现出通用高效性。

中文摘要 AI 辅助

皮层计算具有显著的低维特性:神经群体活动承载的少量潜在变量,调控着单个神经元的高维响应。本研究的目标是实现样本高效性——即能从有限数据中良好解码且参数预算小的模型。标准Transformer层中,前馈模块对每个token应用相同算子;我们提出一种受冯·诺依曼启发的高效计算假设,作为神经解码的替代方案:控制器解码指令,再执行token特定算子;常规实现(如软混合专家)仅融合其输出,而非算子。我们引入冯·诺依曼状态空间Transformer(VN-SST),这是一种带记忆增强的Transformer,其前馈模块为低维指令库:包含共享基础算子与少量学习得到的低维指令,每个token的编码可合成该token实际使用的权重矩阵。该编码从承载的状态空间记忆的低维投影中读取,因此缓慢的潜在轨迹可作为指令指针,模拟低维动态调控皮层计算的方式。在三个运动皮层神经解码基准上,VN-SST的样本效率远高于现代Transformer,可联合预测神经元放电与解码行为。该模型在数据最稀缺的基准上大幅领先,在另外两个基准上也占优,且能让更长上下文带来准确率提升而非下降。我们评估发现,该网络可将大型指令库压缩至每个token仅几位,因此程序容量是控制通道而非准确率杠杆。同一模型在两个用于语言建模(LLMs)的小型文本基准上也具有更高的参数效率,表明这是一种通用机制。

英文摘要

Cortical computation is strikingly low-dimensional: a handful of latent variables, carried in a neural population's activity, steer the higher-dimensional responses of individual neurons. Our aim is sample efficiency-models that decode well from limited data and at small parameter budgets. In a standard Transformer layer, the feed-forward block applies the same operator to every token. We suggest a von-Neumann inspired hypothesis of efficient computation as an alternative for neural decoding: a controller decodes an instruction and then executes a token-specific operator; the usual realization-a soft mixture of experts-only blends their outputs, not operators. We introduce a von-Neumann State-Space Transformer (VN-SST), a memory-augmented Transformer whose feed-forward block is a low-rank instruction bank: a shared base operator plus a small set of learned low-rank instructions, from which a per-token code synthesizes the weight matrix actually used at that token. The code is read from a low- dimensional projection of a carried state-space memory, so a slow latent trajectory acts as an instruction pointer-mirroring how low-dimensional dynamics may route cortical computation. On three motor-cortex neural-decoding benchmarks, VN-SST is far more data-efficient than a modern Transformer, each jointly predicting spikes and decoding behavior. This model wins by a wide margin on the scarcest benchmark, leads on the other two, and turns longer context into rising rather than falling accuracy. We evaluated that the network compresses a large instruction bank to a few bits per token, so program capacity acts as a control channel, not an accuracy lever. The same model is also more parameter-efficient on two small text benchmarks used for language modeling (LLMs), suggesting a generic mechanism.

发表机构

  • BrainCo(脑聚科技)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑