AI 中文总结
本文针对多分区NUMA GPU中LLM内核访问与内存交互问题,通过分析三类LLM内核实现,提出内存追踪与周期级模拟方法,划分操作数共享模式并给出对应优化策略,强调需针对性内核编程与架构支持。
AI 中文摘要
大型语言模型(LLM)工作负载推动多分区GPU成为扩展计算与内存容量的路径,但其非统一内存访问特性及分区间通信会放大竞争、降低局部性,导致内核延迟未达最优。为解决该问题,本文分析了涵盖权重投影、混合专家(mixture-of-experts)及最先进服务引擎的注意力变体的性能关键LLM内核实现,以表征多分区GPU中的数据访问模式。首先,引入内存追踪分析方法推导工作组级别的数据访问与共享行为,再使用周期级模拟器评估其对延迟的局部性影响。借助这些工具,将LLM内核操作数分为三类工作组间共享模式(全局、部分或私有),并表明不同类别所需优化策略不同,从简单的每工作组固定到子组感知的协同调度。研究结果强调,多分区GPU中需要感知放置的内核编程及针对工作与数据局部性的更智能架构支持。
英文摘要
Large language model (LLM) workloads motivate multi-partition GPUs as a path to scaling compute and memory capacity, but their non-uniform memory access characteristics and inter-partition communication can amplify contention and degrade locality, leading to suboptimal kernel latency. To address this, we analyze performance-critical LLM kernel implementations spanning weight projection, mixture-of-experts, and attention variants of state-of-the-art serving engines to present a characterization of data access patterns in multi-partition GPUs. First, we introduce memory trace analysis methodology to derive workgroup-level data access and sharing behavior, then evaluate the locality implications on latency using a cycle-level simulator. Using these tools, we categorize LLM kernel operands into three inter-workgroup sharing patterns (global, partial, or private) and show that the required optimization strategies differ across categories, from simple per-workgroup pinning to subgroup-aware co-scheduling. Our findings highlight the need for placement-aware kernel programming and smarter architectural support for work and data locality in multi-partition GPUs.
Comments12 pages, 6 figures