Mamba 系列状态空间模型核在可编程 CGLA 上的实现
Mamba-Family State-Space Model Kernels on a Programmable CGLA
浏览论文内容
中文总结 AI 辅助
本文针对边缘推理的功耗与数据移动限制,将 Mamba 系列状态空间模型的投影、SSD 和递归更新核映射到可编程 CGLA(IMAX)上,发现投影 GEMV 是解码瓶颈,并指出可编程 CGLA 适合长归约投影核,而 SSD 和解码支持需边界归约与持久权重。
中文摘要 AI 辅助
边缘和嵌入式推理受到功耗和数据移动的限制。Mamba 系列状态空间模型用序列线性递归替代了注意力机制,但其推理路径结合了密集投影、短归约 SSD 核和递归状态更新。本文将这些核组映射到 IMAX(一种可编程 CPU 接地线性阵列,CGLA)上,并从核执行到 token 级集成进行了测量。投影核与 IMAX 的长归约流水线匹配,而 SSD Step-1 受限于短归约和核边界开销。Mamba-130M 的 token 级集成识别出投影 GEMV 是解码瓶颈。这些结果表明,可编程 CGLA 适合长归约投影核,而 SSD 和解码时投影支持需要边界归约和持久权重执行。
英文摘要
Edge and embedded inference is constrained by power and data movement. Mamba-family state-space models replace attention with sequence-linear recurrence, but their inference path combines dense projections, short-reduction SSD kernels, and recurrent-state updates. This paper maps these kernel groups onto IMAX, a programmable CPU-Grounded Linear Array (CGLA), and measures them from kernel execution to token-level integration. Projection kernels match the long-reduction IMAX pipeline, whereas SSD Step-1 is limited by short reductions and kernel-boundary overheads. Mamba-130M token-level integration identifies projection GEMV as the decode bottleneck. These results show that programmable CGLAs fit long-reduction projection kernels, while SSD and decode-time projection support require boundary reduction and persistent-weight execution.
发表机构
- Nara Institute of Science and Technology(奈良先端科学技术大学院大学)
机构由 AI 辅助整理,请以论文原文为准。