Motif-Mamba:融合网络基序改进Mamba的长程序列建模模型
Motif-Mamba: network motif improved mamba for long-range sequence modeling
- Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出融合网络基序的结构化状态空间模型Motif-Mamba,改进Mamba的对角状态转移限制,在长序列建模任务上性能优于原Mamba,为长程序列建模提供有效结构先验。
AI中文摘要:
高效长序列建模仍是大语言模型的核心挑战,因为自注意力机制的计算复杂度随序列长度呈二次方增长。Mamba通过选择性状态空间循环提供了线性时间复杂度的替代方案,但其以对角为主的状态转移限制了状态维度间的显式交互。我们提出Motif-Mamba,一种结构化状态空间模型,它为Mamba增加了基序约束的低秩循环通路。受三节点网络基序的动力学启发,该通路将隐状态投影到紧凑的动力学子空间,施加基序引导的交互,并将所得动力学映射回原始状态空间。此设计在保留Mamba线性时间循环结构的同时增强了跨维度通信。在长序列外推、语言建模基准及脑机接口解码任务上的实验显示,其性能一致优于Mamba主干模型,表明基序引导的低秩动力学为长程序列建模提供了有效的结构先验。
英文摘要:
Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length. Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions among state dimensions. We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent pathway. Inspired by the dynamics of three-node network motifs, the proposed pathway projects hidden states into a compact dynamical subspace, imposes motif-guided interactions, and maps the resulting dynamics back to the original state space. This design enhances cross-dimensional communication while preserving the linear-time recurrent structure of Mamba. Experiments on long-sequence extrapolation, language modeling benchmarks, and brain--computer interface decoding show consistent improvements over Mamba backbones, suggesting that motif-guided low-rank dynamics provide an effective structural prior for long-range sequence modeling.