AI 中文总结
本文提出PSC方法,通过解析器栈分类实现高效语法约束解码,其计算掩码速度远快于基线,可提升LLM的端到端吞吐量。
AI 中文摘要
大型语言模型(LLM)被广泛用于生成源代码或JSON等结构化输出。语法约束解码(GCD)可通过屏蔽违反上下文无关文法规则的token,保证生成输出的句法有效性。然而,现有GCD方法的在线计算开销通常随词汇量呈线性缩放,限制了LLM的吞吐量,尤其对于词汇量大的模型。为解决该问题,本文提出一种名为PSC的新型语法约束解码方法。PSC通过预处理将所有词汇token的接受条件合并为解析器栈的单一分类器,每个解码步骤仅需检查一次解析器栈即可计算完整的词汇掩码,时间复杂度与词汇量无关。实验表明,在复杂编程语言文法下,PSC计算掩码的速度比基线快700倍,在符合模式的JSON场景下快30倍;采用PSC的LLM端到端吞吐量接近无约束解码的水平。本文还分析了预处理提供者和解码用户的预处理开销,并提供了盈亏平衡点分析,以帮助用户决定是否自行进行预处理。
英文摘要
LLMs are widely used to generate structured output like source code or JSON. Grammar-constrained decoding (GCD) can guarantee the syntactic validity of the generated output, by masking out tokens that violate rules specified by a context-free grammar. However, the online computational overhead of existing GCD methods, with latency typically scaling linearly with vocabulary size, limits the throughput of LLMs, especially for models with large vocabularies. To address this issue, we propose PSC, a novel grammar-constrained decoding method. By combining acceptance conditions of all vocabulary tokens into a single classifier of the parser stack during preprocessing, PSC can compute the complete vocabulary mask by checking the parser stack exactly once per decoding step, with time complexity independent of the vocabulary size. Experiments show that PSC computes masks up to 700$\times$ faster than baselines on complex programming language grammars, and up to 30$\times$ faster for schema-conformant JSON; end-to-end LLM throughput with PSC approaches that of unconstrained decoding. We analyze the preprocessing overhead for preprocessing providers and decoding users, and provide a break-even point analysis to help users decide whether to do preprocessing by themselves.
Commentsaccepted by ISSTA 2026