arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06989cs.AR

重新思考NPU-PIM系统的统一内存:面向大语言模型动态推理的双视图内存

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM

Shixin Zhao, Lian Liu, Tianhua Han, Mengdi Wang, Yinhe Han, Ying Wang

AI总结:

针对NPU-PIM系统现有统一内存无法适配LLM动态推理导致的带宽浪费问题,提出双视图内存方案PFM,可将端到端吞吐量提升最高达2.32倍。

AI中文摘要:

结合神经处理单元(NPU)与内存处理单元(PIM)的异构架构正被越来越多地用于加速大语言模型(LLM)推理。现有研究聚焦于构建统一内存,使NPU与PIM无需重复即可共享数据,但这些设计默认每个张量绑定固定执行设备,因此依赖静态、偏向设备的数据映射。我们发现该假设在现代LLM工作负载中不成立:由于阶段变化(如预填充与解码)及MoE路由等动态行为,同一张量的最优执行设备可在运行时改变。在这类动态执行下,偏向设备的映射会与访问模式不匹配,导致带宽大量未利用与性能损失。本文提出PFM(PIM-as-Flexible-Memory),一种双视图内存系统,将物理数据布局与访问器可见的逻辑视图解耦。PFM以联合优化的物理布局存储数据,向NPU与PIM暴露不同的逻辑解释,支持跨设备高效访问且无需数据复制或重布局。我们进一步设计感知访问器的地址转换与运行时调度机制,以支持LLM工作负载波动及最优执行设备动态变化时的动态执行。对多款LLM的评估显示,PFM可将端到端吞吐量提升最高达2.32倍,证明其作为NPU-PIM系统统一内存管理方案的有效性与广泛适用性。

英文摘要:

Heterogeneous architectures that combine neural processing unit (NPU) and processing-in-memory (PIM) are increasingly adopted to accelerate LLM inference. Prior work focuses on building a unified memory that allows NPUs and PIM to share data without duplication. However, these designs implicitly assume that each tensor is bound to a fixed execution device, and therefore rely on static, device-biased data mappings. We observe that this assumption does not hold in modern LLM workloads. Due to phase changes (e.g., prefill vs. decode) and dynamic behaviors such as MoE routing, the optimal execution device for the same tensor can change at runtime. Under such dynamic execution, device-biased mappings become mismatched to access patterns, leading to substantial bandwidth underutilization and performance loss. This paper presents PFM (PIM-as-Flexible-Memory), a dual-view memory system that decouples physical data layout from accessor-visible logical views. PFM stores data in a jointly optimized physical layout and exposes different logical interpretations to NPUs and PIM, enabling efficient access across devices without data duplication or relayout. We further design accessor-aware address translation and runtime scheduling mechanisms to support dynamic execution when LLM workloads fluctuate and the optimal execution device dynamically changes. Our evaluation across LLMs shows that PFM improves end-to-end throughput by up to 2.32$\times$, demonstrating its effectiveness and broad applicability as a unified memory management solution for NPU-PIM systems.

补充信息

↑