arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlashTrie:用于生成式检索的GPU加速约束束搜索

FlashTrie: A GPU-Accelerated Constrained Beam Search for Generative Retrieval

Dakshitha Anandakumar, Anurag Mukkara, Wenxiang Hu, Jiusheng Chen, M Akash Kumar, Ting Ye, Qiang Lou, Jian Jiao

arXiv 2607.10044首次发表:更新:

发表机构

Microsoft; Nvidia(微软; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对生成式检索中约束解码的瓶颈问题,提出FlashTrie方法。通过优化GPU上的约束束搜索,采用整数感知简洁trie布局和协作CUDA内核等技术,显著降低解码延迟、提高吞吐量,在实验中取得良好效果并提升了收入。

AI 中文摘要

约束解码在生成式检索中至关重要,直接从查询生成的文档标识符必须与预定义的有效ID库完全匹配。大规模时,通常使用带有束搜索的trie进行约束解码,但大多数实现运行在CPU上。随着束宽度增加,有限的并行性使trie遍历和候选验证成为服务瓶颈。我们提出FlashTrie,通过在GPU上优化约束束搜索来解决此限制。它引入整数感知简洁trie布局,使用位压缩减少内存占用,同时将完整索引保存在GPU高带宽内存中以减少内存停顿;还引入协作CUDA内核,完全在设备上执行束扩展、验证和修剪,无需主机逐步骤编排。它进一步用GPU感知并行原语取代CPU风格的不规则查找和堆维护,提高线程利用率并减少分歧。这些设计显著降低解码延迟并提高吞吐量,同时保持检索质量。在包含8亿关键词且束宽度高达1000的库上,FlashTrie将trie搜索延迟降低到3毫秒以下,比高度优化的多线程CPU基线实现高达24倍的加速。这些改进使FlashTrie在诸如赞助搜索等延迟关键应用中能够将束大小扩展多达5倍。在一个流行商业搜索引擎上的大规模在线A/B实验中,它带来了统计学上显著的0.71%的收入提升,实现了以前仅离线可行规模的实时约束解码。FlashTrie代码将在评审过程后公开发布。

英文摘要

Constrained decoding is essential in generative retrieval, where document identifiers generated directly from a query must exactly match a predefined library of valid IDs. At scale, decoding is often constrained using a trie with beam search but most implementations run on CPU. Limited parallelism then makes trie traversal and candidate validation a serving bottleneck as beam width grows. We present FlashTrie, which addresses this limitation by optimizing constrained beam search on GPUs. It introduces an integer-aware succinct trie layout that uses bit compression to reduce memory footprint while keeping the full index in GPU high-bandwidth memory reducing memory stalls, and a cooperative CUDA kernel that performs beam expansion, validation, and pruning entirely on-device without per-step host orchestration. It further replaces CPU-style irregular lookup and heap maintenance with GPU-aware parallel primitives, improving warp utilization and reducing divergence. Together, these designs significantly reduce decoding latency and increase throughput while preserving retrieval quality. On a library of 800M keywords with beam widths up to 1000, FlashTrie reduces trie-search latency to under 3 ms, achieving up to 24x speedup over a highly optimized multi-threaded CPU baseline. These improvements enable FlashTrie to scale beam sizes by up to 5x in latency-critical applications such as sponsored search. In a large-scale online A/B experiment on a popular commercial search engine, it delivers a statistically significant +0.71% revenue lift, enabling real-time constrained decoding at a scale previously feasible only offline. The FlashTrie code will be publicly released after the review process.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑