arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FastTPS:一种针对人工智能加速器的大语言模型令牌阶段的优化方法

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators

Wenzong Yang, Danyang Zhang, Kun Cao, Tejus Siddagangaiah, Rajeev Patwari, Zhanxing Pu, Siyin Kong, Zijiang Yang, Hao Zhu, Varun Sharma, Yue Gao, Tianping Li, Fan Yang, Jicheng Chen, Yushan Chen, Fennian Zhao, Aaron Ng, Elliott Delaye, Ashish Sirasao, Sudip Nag

arXiv 2607.11211首次发表:更新:

发表机构

Anonymous Institution(匿名机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对LLMs推理令牌阶段低并行性问题,提出FastTPS方法,含无重新加载KV缓存串联、优化“RoPE”注意力、高度融合MLP及细粒度流水线调度,显著缓解内存瓶颈,提升推理速度与内存带宽利用率。

AI 中文摘要

大语言模型(LLMs)的普及对有效推理的需求不断升级。在仅解码器的LLMs推理的令牌阶段,由于令牌的顺序处理,固有的低并行性导致人工智能(AI)加速器上计算单元的吞吐量降低和利用率次优,特别是处理长序列输入时内存开销大。许多方法有数值偏差。本文提出FastTPS,一种用于在通用AI加速器上加速LLM推理令牌阶段的高性能低精度损失方法,包括三个关键组件:支持AI加速器的无重新加载KV缓存串联、基于平铺优化FLAT的高效高精度“RoPE”注意力、具有细粒度流水线调度的高度融合MLP。结果表明FastTPS显著缓解令牌阶段内存瓶颈,在AMD Ryzen AI 300系列NPU上,BF16精度下比无融合提高6倍速度,在Phi3-mini-4k-instruct推理中维持93%的峰值内存带宽利用率。

英文摘要

The popularity of large language models (LLMs) escalates an ongoing demand for effective inference. However, due to the sequential processing of tokens during the token phase in decoder-only LLMs inference, the inherent low parallelism leads to reduced throughput and suboptimal utilization of the computing units on artificial intelligence (AI) accelerators, particularly when handling long-sequence inputs that impose significant memory overhead. Recently, many reported methods have been developed as potential solutions, since they emerge with numeric deviation. This paper presents FastTPS, a high performance and low-precision loss method for accelerating the token-phase in LLM inference on general AI accelerators which includes three key components: (1) AI accelerator-enabled reloading-free KV Cache concatenation which decreases memory access overhead as well as enables full fusion of Attention, (2) high-efficiency and high-accuracy 'RoPE' attention based on the tiling optimized FLAT, and (3) highly-fused MLP with fine-grain pipeline scheduling. Our results confirm that FastTPS significantly alleviates memory bottlenecks in the token phase, delivering a 6x speed improvement (compared to none-fusion) on an AMD Ryzen AI 300 series NPU with BF16 precision while sustaining 93% peak memory bandwidth utilization during Phi3-mini-4k-instruct inference.

Comments16 pages, 8 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑