发表机构
Anonymous Institution(匿名机构)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLMs推理令牌阶段低并行性问题,提出FastTPS方法,含无重新加载KV缓存串联、优化“RoPE”注意力、高度融合MLP及细粒度流水线调度,显著缓解内存瓶颈,提升推理速度与内存带宽利用率。
AI 中文摘要
大语言模型(LLMs)的普及对有效推理的需求不断升级。在仅解码器的LLMs推理的令牌阶段,由于令牌的顺序处理,固有的低并行性导致人工智能(AI)加速器上计算单元的吞吐量降低和利用率次优,特别是处理长序列输入时内存开销大。许多方法有数值偏差。本文提出FastTPS,一种用于在通用AI加速器上加速LLM推理令牌阶段的高性能低精度损失方法,包括三个关键组件:支持AI加速器的无重新加载KV缓存串联、基于平铺优化FLAT的高效高精度“RoPE”注意力、具有细粒度流水线调度的高度融合MLP。结果表明FastTPS显著缓解令牌阶段内存瓶颈,在AMD Ryzen AI 300系列NPU上,BF16精度下比无融合提高6倍速度,在Phi3-mini-4k-instruct推理中维持93%的峰值内存带宽利用率。
英文摘要
The popularity of large language models (LLMs) escalates an ongoing demand for effective inference. However, due to the sequential processing of tokens during the token phase in decoder-only LLMs inference, the inherent low parallelism leads to reduced throughput and suboptimal utilization of the computing units on artificial intelligence (AI) accelerators, particularly when handling long-sequence inputs that impose significant memory overhead. Recently, many reported methods have been developed as potential solutions, since they emerge with numeric deviation. This paper presents FastTPS, a high performance and low-precision loss method for accelerating the token-phase in LLM inference on general AI accelerators which includes three key components: (1) AI accelerator-enabled reloading-free KV Cache concatenation which decreases memory access overhead as well as enables full fusion of Attention, (2) high-efficiency and high-accuracy 'RoPE' attention based on the tiling optimized FLAT, and (3) highly-fused MLP with fine-grain pipeline scheduling. Our results confirm that FastTPS significantly alleviates memory bottlenecks in the token phase, delivering a 6x speed improvement (compared to none-fusion) on an AMD Ryzen AI 300 series NPU with BF16 precision while sustaining 93% peak memory bandwidth utilization during Phi3-mini-4k-instruct inference.
Comments16 pages, 8 figures, 7 tables