arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于大有限集上约束解码的Trie自动机

Trie Automata for Constrained Decoding over Large Finite Sets

Xingzi Xu, Karim Bouyarmane

arXiv 2608.12574首次发表:更新:

发表机构

Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型语言模型有限集约束解码的速度瓶颈,提出Trie自动机机制,通过预计算token掩码实现高效解码,在批量服务中吞吐量较XGrammar提升29倍,且编译与计算效率显著提升。

AI 中文摘要

大型语言模型越来越需要生成符合预定义模式的结构化输出,其中一个常见约束是从有限的有效字符串集中进行选择。当前的约束解码系统通过通用语法编译来处理这一问题,当有效值数量增长到数千时,会出现严重的速度瓶颈,即基数障碍。我们引入了Trie自动机,这是一种利用有限集结构(共享前缀、有限深度、已知基数)的专用机制,通过Aho-Corasick多模式匹配来预计算每个节点的token掩码。与vLLM和SGLang中的主要后端之一XGrammar相比,Trie自动机的每一步有效token计算速度快7倍(0.65微秒对5.8微秒),且在K≥300时编译速度快2至6.5倍。由于预计算的掩码支持绕过引导解码流水线的无状态服务路径,这一优势在批量服务中会被放大:当批量大小为256时,端到端vLLM吞吐量达到219请求/秒,而XGrammar为7.5请求/秒,提升了29倍,这29倍的提升结合了算法加速和仅预计算掩码才能实现的集成路径节省。在7个tokenizer系列(词汇量为32K至262K)中,Trie自动机在K=10000时仍保持低于100毫秒的编译时间,且每一步的成本与集合大小无关,同时保证100%的输出有效性。

英文摘要

Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar compilation, which becomes prohibitively slow as the number of valid values grows into the thousands, a cardinality wall. We introduce the trie automaton, a specialized mechanism that exploits finite-set structure (shared prefixes, bounded depth, known cardinality) via Aho-Corasick multi-pattern matching to precompute per-node token masks. The trie achieves 7X faster per-step valid-token computation (0.65 us vs. 5.8 us) compared to XGrammar, one of the primary backends in vLLM and SGLang, and 2--6.5X faster compilation at K >= 300. Because precomputed masks enable a stateless serving path that bypasses the guided decoding pipeline, this advantage compounds in batch serving: end-to-end vLLM throughput reaches 219 req/s vs. XGrammar's 7.5 req/s at batch size 256 (29X). The 29X combines the algorithmic speedup with integration-path savings that only precomputed masks can unlock. Across seven tokenizer families (32K--262K vocabulary), the trie maintains sub-100ms compilation up to K = 10,000 and flat per-step cost regardless of set size, while guaranteeing 100% output validity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑