NOVA:面向注意力-状态空间模型-混合专家混合大语言模型推理的近内存处理技术-架构协同设计
NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference
浏览论文内容
中文总结 AI 辅助
NOVA是面向注意力-SSM-MoE混合LLM推理的近内存处理技术-架构协同设计方案,通过4F² VCT DRAM与两层NMP架构,在低面积开销下实现远超GPU的推理性能与能效。
中文摘要 AI 辅助
混合大语言模型(LLM)通过交替使用分组查询注意力(GQA)、状态空间模型(SSM)和混合专家(MoE)层实现快速演进,这给近内存处理(NMP)架构带来两大核心挑战。一是技术壁垒:传统6F² DRAM单元在10nm级节点已接近物理缩放极限,难以满足拥有数百个专家的MoE模型对内存容量的需求。二是架构壁垒:现有NMP设计仅针对窄算术强度(Op/B)范围,无法高效支持混合LLM的异构计算特性,这类模型涵盖内存密集型SSM层、计算密集型GQA层,且不同专家间Op/B差异巨大。本文提出NOVA,一款克服上述两大壁垒的技术-架构协同设计NMP系统。技术层面,NOVA将4F²垂直沟道晶体管(VCT)DRAM单元与单元周边(POC)结构结合,在相同面积下实现约2倍于传统6F² DRAM的内存密度,可延续至10nm以下节点的缩放。架构层面,NOVA将POC周边芯片(peri-die)重新用于部署处理单元,形成两层NMP架构:第一层(peri-die NMP)面向中低Op/B操作,第二层(base-die NMP)面向中高Op/B操作。层间并行执行支持混合LLM的多样计算模式,最大化推理性能。在Nemotron3-Nano、Nemotron3-Super、Falcon-H1R和Qwen3等当前顶尖混合及MoE LLM上的评估显示,与GPU基准相比,NOVA平均实现4.5倍吞吐量提升、69.8%端到端延迟降低、5倍能效提升,仅产生3.9%的面积开销且无内存容量损失。
英文摘要
The rapid evolution of hybrid large language models (LLMs), which interleave grouped-query-attention (GQA), state-space model (SSM), and Mixture-of-Experts (MoE) layers, introduces two fundamental challenges for near-memory processing (NMP) architectures. First, the Technology Wall: the conventional 6F^2 DRAM cell is approaching its physical scaling limits at 10nm-class nodes, making it difficult to meet the memory capacity demands of MoE models with hundreds of experts. Second, the Architecture Wall: existing NMP designs target narrow arithmetic intensity (Op/B) ranges and cannot efficiently support the heterogeneous compute characteristics of hybrid LLMs, spanning memory-bound SSM layers, compute-intensive GQA layers, and large Op/B variations across experts. We propose NOVA, a technology-architecture co-designed NMP system that overcomes both walls. On the technology side, NOVA combines a 4F^2 vertical channel transistor (VCT) DRAM cell with a peri-over-cell (POC) structure to achieve approximately 2x memory density at iso-area over conventional 6F^2-based DRAM, enabling continued scaling into sub-10nm nodes. On the architecture side, NOVA repurposes the POC peripheral-die (peri-die) to host processing units, forming a 2-tier NMP architecture: Tier-1 (peri-die NMP) for low-to-mid Op/B operations, and Tier-2 (base-die NMP) for mid-to-high Op/B operations. Parallel execution across tiers supports diverse compute patterns for hybrid LLMs, maximizing inference performance. Evaluated on state-of-the-art hybrid and MoE LLMs including Nemotron3-Nano, Nemotron3-Super, Falcon-H1R, and Qwen3, NOVA achieves on average 4.5x higher throughput, 69.8% lower end-to-end latency, and 5x better energy efficiency over a GPU baseline, with only 3.9% area overhead and no loss in memory capacity.