arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10137cs.CLcs.LG

解析器已掌握:约束解码中的轻量级偏差校正

The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding

Işıl Özgü, Yaoxuan Wu, Guy Van den Broeck, Miryung Kim

AI总结:

该研究针对语法约束解码中刚性屏蔽扭曲模型分布的问题,利用增量解析的语法状态提出轻量级离线 logit 校正,在几乎不增加开销的情况下优于基线方法,兼顾了语言模型的概率完整性与语法一致性。

AI中文摘要:

语法约束解码(Grammar Constrained Decoding, GCD)迫使语言模型(Language Models, LMs)在每一步通过屏蔽不符合语法的 token 来生成语法有效的输出,但刚性屏蔽会扭曲模型的基础概率分布,常使生成结果偏向语法有效但次优的输出。虽然在线采样可恢复该分布,却需要计算成本高昂的迭代重采样,导致现有方法需在输出质量与推理延迟间妥协。本文核心洞见是:增量解析过程中固有维护的内部解析器和词法分析器状态已编码了未来语法有效性,这正是恢复 LM 真实分布所需的信息。我们提出一种轻量级、离线训练的 logit 校正,其条件为该语法与词法状态及候选下一个 token。由于这些状态是屏蔽所需的增量解析必要部分,提取它们的开销可忽略不计,且完全不改动基础 LM 的权重。在多种语法上,该校正大幅缩小了屏蔽分布与 LM 真实分布的差距,始终优于屏蔽和在线采样;即便仅依赖候选下一个 token 的最轻量变体,也仍能达到或超过两种基线——下一个 token 本身带有隐式前瞻,类似解析器常用前瞻 token 解决歧义决策的方式。通过恢复屏蔽所去除的概率质量,该校正协调了 LM 的概率完整性与语法一致性。

英文摘要:

Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. However, because masking only checks whether each token is valid so far, the resulting distribution over complete outputs diverges from the LM's own distribution conditioned on the grammar, biasing generation toward valid but suboptimal outputs. Online sampling can restore this distribution, but only through costly iterative resampling. Our key insight is that the parser and lexer states that GCD tools already maintain carry a strong signal about future grammatical validity. We introduce SHIM, a lightweight, offline-trained correction of the LM's next-token probabilities, conditioned on this syntactic and lexical state together with candidate next tokens. Since GCD tools already compute these states, SHIM leaves the LM itself untouched. Across bit-vector and text-to-SQL grammars, this correction substantially narrows the gap to the LM's grammar-conditioned distribution compared to masking and online sampling, while running at nearly masking's speed. Even a variant that sees only the next token can improve on both baselines, making SHIM usable with GCD tools that do not expose their parser state.

补充信息

↑