arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00027cs.AI

Motif-Mamba:融合网络基序改进Mamba的长程序列建模模型

Motif-Mamba: network motif improved mamba for long-range sequence modeling

  • Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心)

机构由 AI 辅助整理,请以论文原文为准。

Chonghe Hao, Yue Sun, Jian Zhang, Yansong Wang, Wangzi Yao, Yunjie Yao, Tielin Zhang

AI总结:

该研究提出融合网络基序的结构化状态空间模型Motif-Mamba,改进Mamba的对角状态转移限制,在长序列建模任务上性能优于原Mamba,为长程序列建模提供有效结构先验。

AI中文摘要:

高效长序列建模仍是大语言模型的核心挑战,因为自注意力机制的计算复杂度随序列长度呈二次方增长。Mamba通过选择性状态空间循环提供了线性时间复杂度的替代方案,但其以对角为主的状态转移限制了状态维度间的显式交互。我们提出Motif-Mamba,一种结构化状态空间模型,它为Mamba增加了基序约束的低秩循环通路。受三节点网络基序的动力学启发,该通路将隐状态投影到紧凑的动力学子空间,施加基序引导的交互,并将所得动力学映射回原始状态空间。此设计在保留Mamba线性时间循环结构的同时增强了跨维度通信。在长序列外推、语言建模基准及脑机接口解码任务上的实验显示,其性能一致优于Mamba主干模型,表明基序引导的低秩动力学为长程序列建模提供了有效的结构先验。

英文摘要:

Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length. Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions among state dimensions. We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent pathway. Inspired by the dynamics of three-node network motifs, the proposed pathway projects hidden states into a compact dynamical subspace, imposes motif-guided interactions, and maps the resulting dynamics back to the original state space. This design enhances cross-dimensional communication while preserving the linear-time recurrent structure of Mamba. Experiments on long-sequence extrapolation, language modeling benchmarks, and brain--computer interface decoding show consistent improvements over Mamba backbones, suggesting that motif-guided low-rank dynamics provide an effective structural prior for long-range sequence modeling.

补充信息

↑