HEMERA:一种用于边缘受限状态空间对偶模型推理的具有递归数据流的异构内存中心加速器
HEMERA: A Heterogeneous Memory-Centric Accelerator with Recursive Dataflow for Edge-Constrained State-Space-Duality Models Inference
浏览论文内容
中文总结 AI 辅助
针对结构化状态空间模型推理中SSD存在系统级开销的问题,本文提出异构内存中心加速器HEMERA,将SSD计算重新表述为流递归数据流以避免二次中间存储,在不同规模Mamba - 2模型上加速效果显著,展示了边缘约束下高效部署的潜力。
中文摘要 AI 辅助
结构化状态空间模型(SSMs),如Mamba,能够以线性时间复杂度进行高效的长序列建模。近期实现通过结构化状态空间对偶性(SSD)实现此能力,它将递归状态演化转换为矩阵形式计算。然而,SSD引入大量系统级开销。尽管先前加速器通过优化数据流或内存计算技术减轻了这些开销,但仍存在问题。本文提出HEMERA,一种用于高效Mamba - 2推理的异构内存中心加速器。它将SSD计算重新表述为代数等效的流递归数据流,避免二次中间存储。在不同规模的Mamba - 2模型上,HEMERA相对于NVIDIA A100上的官方优化融合Mamba - 2内核,平均延迟加速1.4倍至3.6倍,能源效率提高12.2倍至27.0倍,在长序列推理中进一步降低平均SSD相关执行时间比例至14.12%,展示了其在边缘约束下高效部署的潜力。
英文摘要
Structured State Space Models (SSMs), such as Mamba, enable efficient long-sequence modeling with linear time complexity. Recent implementations realize this capability through Structured State Space Duality (SSD), which transforms recursive state evolution into matrix-form computations. However, SSD introduces substantial system-level overheads, including quadratic intermediate materialization, irregular data movement, and prefix-dependent execution, leading to excessive memory traffic and bandwidth demand on conventional architectures. Although prior accelerators mitigate these overheads through optimized dataflows or compute-in-memory techniques, they largely retain matrix-oriented SSD execution and cannot simultaneously avoid quadratic intermediate storage and efficiently map dependency-bound state propagation. This paper presents HEMERA, a heterogeneous memory-centric accelerator for efficient Mamba-2 inference. Rather than directly executing the matrix-form SSD computation, HEMERA reformulates it into an algebraically equivalent streaming-recursive dataflow that avoids quadratic intermediate storage while preserving the original computation. The resulting heterogeneous execution paradigm maps dense linear operations onto in-memory computing units and recursive state updates onto a dedicated streaming engine. Across Mamba-2 models ranging from 130M to 2.8B, HEMERA achieves average latency speedups of 1.4x-3.6x and energy-efficiency improvements of 12.2x-27.0x over the official optimized fused Mamba-2 kernel on NVIDIA A100. It further reduces the average SSD-related execution-time ratio across model scales to 14.12% during long-sequence inference, demonstrating its potential for efficient deployment under edge constraints.