MeshKV:面向可扩展Transformer解码加速器的片上网络KV缓存结构
MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators
浏览论文内容
中文总结 AI 辅助
MeshKV提出一种基于轻量级NoC的KV缓存结构,通过条带化、多播和流水线重叠技术,在FPGA上实现最高58%的互连流量削减和2.1倍带宽利用率提升。
中文摘要 AI 辅助
自回归Transformer解码在瓦片式加速器上受到不规则键值(KV)缓存移动的制约。先前的压缩和DRAM放置系统仍将流量集中在集中式存储路径上,这成为长上下文服务的瓶颈。我们提出MeshKV,一种KV缓存结构,它将块作为分组化流在轻量级片上网络(NoC)上移动。它协同设计了(i) TaKV仿射条带化以分散归属并减少热点负载,(ii) 带有验证重复抑制的Mare多播,以及(iii) Pad,它在信用对齐的FIFO之后重叠预取、瓦片乘法和流式softmax。这些设计共同将对分背压转化为有用的KV传输。在我们搭载LLaMA-2-7B和Mistral-7B、上下文长度8K-32K的8x8 FPGA实现上,MeshKV将互连流量减少高达58%,将KV带宽利用率提升2.1倍,并提供高达1.9倍的多流吞吐量。
英文摘要
Autoregressive transformer decoding is constrained by irregular key-value (KV) cache movement on tiled accelerators. Prior compression and DRAM-placement systems still concentrate traffic on centralized memory paths that bottleneck long-context serving. We present MeshKV, a KV cache fabric that moves blocks as packetized flows over a lightweight NoC. It co-designs (i) TaKV affine striping to spread homes and cut hotspot load, (ii) Mare multicast with verified duplicate suppression, and (iii) Pad, which overlaps prefetch, tile multiply, and streaming softmax behind credit-aligned FIFOs. Together they convert bisection back-pressure into useful KV transfer. On our 8x8 FPGA implementation with LLaMA-2-7B and Mistral-7B at 8K-32K, MeshKV reduces interconnect traffic by up to 58%, improves KV bandwidth utilization by 2.1x, and delivers up to 1.9x multi-stream throughput.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。