arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03555cs.AR

面向基于检索的稀疏注意力的通用近内存处理异构大语言模型服务

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention

Hyungkyu Ham, Junhyeong Bae, Seungheon Lee, Myeongjae Jeon, Gwangsun Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对前沿LLM的基于检索的稀疏注意力需求,提出异构解码服务系统KARAT,结合PNM与OFMS、CMR优化,大幅提升吞吐量与每TDP性能。

中文摘要 AI 辅助

本文提出一种异构解码阶段服务系统,受近期前沿大语言模型(LLM)采用基于检索的稀疏注意力来服务百万token上下文的驱动,该系统将KV缓存移出GPU内存。它按操作类型划分解码步骤:GPU节点存储模型权重并执行投影与混合专家(MoE)层,而近内存处理(PNM)节点存储KV缓存和索引键,并执行所有读取这些数据的操作。我们首先证明,现有内存内处理(PIM)和PNM设计背后的假设已不适用于这些操作,并推导了此类节点的四项设计要求。基于这些要求,我们提出KARAT(面向基于检索的注意力的KV缓存驻留加速器),这是一种通用PNM设计,是满足全部四项要求的设计方案。KARAT设备结合了大容量低功耗双倍数据率内存(LPDDR)与适配检索索引器规模的通用计算能力,其运行强度超出为低强度通用矩阵向量乘法(GEMV)设计的PIM/PNM设计的目标,同时可支持固定功能单元无法适配的各类稀疏注意力算法,这类算法会随技术发展而演进。为减少两种设备类型在微批次间交替时产生的流水线气泡,我们进一步提出机会型细粒度微批次调度(OFMS),该方法将专家全对全通信隐藏在另一微批次的GEMM操作之后;还提出感知上下文长度的微批次再平衡(CMR),该方法可均衡各微批次的token数量,即便上下文长度存在差异。在三种最先进模型和真实智能体(agentic)轨迹上,我们提出的系统在服务水平目标下,相较于仅使用GPU的基线,每TDP(热设计功耗)的吞吐量提升了2.09至6.13倍,且运行无需训练的稀疏注意力方法时,性能提升达1.36至3.21倍。

英文摘要

This paper presents a heterogeneous decode-phase serving system that relocates the KV cache out of GPU memory, motivated by the retrieval-based sparse attention that recent frontier LLMs adopt to serve million-token contexts. It partitions a decode step by operation type: GPU nodes hold the model weights and execute the projections and MoE layers, while processing-near-memory (PNM) nodes hold the KV cache and index keys and execute every operation that reads them. We first show that the assumptions behind prior PIM and PNM designs no longer hold for these operations, and derive four design requirements for such a node. From these requirements, we propose KARAT (KV-cache-resident Accelerator for Retrieval-based ATtention), a general-purpose PNM design that is the design point meeting all four. A KARAT device combines large LPDDR capacity with general-purpose compute sized for the retrieval indexer, serving an operational intensity beyond what PIM/PNM designs built for low-intensity GEMV target while accommodating diverse sparse attention algorithms that fixed-function units cannot support as they evolve. To reduce pipeline bubbles as the two device types alternate between micro-batches, we further propose opportunistic, fine-grained micro-batch scheduling (OFMS), which hides expert all-to-all behind the other micro-batch's GEMMs, and context-length-aware micro-batch rebalancing (CMR), which equalizes their token counts despite the variance in context length. Across three state-of-the-art models and real agentic traces, our proposed system improves throughput per TDP under a service-level objective by 2.09-6.13x over a GPU-only baseline and runs training-free sparse attention methods with 1.36-3.21x improvements.

↑