LLBPE:基于链表的GPU并行BPE分词器
LLBPE: Linked-List Based GPU-Parallel BPE Tokenizer
查看机构详情
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对GPU并行BPE分词器受数据移动瓶颈限制的问题,提出基于数组链表的LLBPE,融合排名查找、最小值选择与合并于单内核,吞吐量较最佳GPU实现提升5.2倍,较优化CPU实现提升24.6倍。
中文摘要 AI 辅助
每一次大语言模型(LLM)推理都始于分词,即将原始输入字节转换为模型所消费的离散词元序列。对于文本,该步骤通常使用字节对编码(BPE)实现,这是一种最初为数据压缩而引入的算法。BPE传统上在CPU上运行并经过大量优化,但近期工作已将其移至GPU以获得更高吞吐量。我们发现这些GPU实现的瓶颈不在于计算,而在于数据移动。我们开发了LLBPE,它将词元序列表示为基于数组的链表,使得每次合并简化为常数时间的指针更新。此外,LLBPE将排名查找、最小值选择和合并融合到单个内核中,以消除冗余的哈希表查询。LLBPE相比现有最佳GPU实现吞吐量最高提升5.2倍,相比优化后的CPU实现提升24.6倍,代价是分词输出存在轻微差异。
英文摘要
Every LLM inference begins with tokenization, which converts raw input bytes into the discrete token sequence the model consumes. For text, this step is often implemented using Byte Pair Encoding (BPE), an algorithm originally introduced for data compression. BPE has traditionally run on the CPU with extensive optimization, but recent work has moved it to the GPU for higher throughput. We show that these GPU implementations are bottlenecked not by computation but by data movement. We develop LLBPE that represents the token sequence as an array-based linked list so that each merge reduces to a constant- time pointer update. Furthermore, LLBPE fuses rank lookup, minimum selection, and merging into a single kernel to eliminate redundant hash map queries. LLBPE achieves up to 5.2x higher throughput than the best existing GPU implementation and 24.6x over optimized CPU implementations, at the cost of minor discrepancies in tokenized output.