arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GrammarRL:通过强化学习实现有效的语法约束解码

GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning

Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov\`ı

arXiv 2609.39869首次发表:更新:

发表机构

University of Catania; ISTC - National Research Council, Italy; University of Bologna(卡塔尼亚大学; 意大利国家研究委员会科学技术研究所; 博洛尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GrammarRL提出无标签强化学习方法,通过直接与反向奖励优化语言模型适应语法约束,在多项任务上平均提升9.8分,最高达22.8 BLEU,且保持贪心解码成本。

AI 中文摘要

语法约束生成保证了句法有效性,但当模型的首选输出与所施加的语法对齐不佳时,可能会显著降低语义质量。当提示词规定不充分或模型的指令遵循能力有限时,这种权衡尤为严重。束搜索可以通过探索多个有效序列来部分缓解这些失败,但其计算成本随束宽增长,而序列级概率仅是语义质量的不完美代理。我们提出GrammarRL,一种无标签的强化学习方法,无需标注数据即可使语言模型适应语法约束。GrammarRL使用两个源自其自身似然度的互补自监督奖励来优化模型:直接奖励,衡量给定输入下约束输出的可能性;以及反向奖励,衡量从生成的输出中重建输入的效果。我们使用Reinforce Leave-One-Out(RLOO)目标在语法约束的rollout组上优化这些奖励,并辅以top-1束搜索假设,同时向冻结的基础模型进行正则化。我们使用参数规模从1B到8B的Llama模型,在手语词汇翻译、层次文本分类和命名实体识别任务上评估GrammarRL。GrammarRL始终优于约束贪心解码,平均提升9.8分,最高提升22.8 BLEU。它在三个任务中的两个上达到或超过束搜索性能,同时保持贪心解码的推理成本。消融实验进一步表明,这两个奖励是互补的:单独使用任一奖励可能不如未训练基线,而它们的组合则持续优于基线。

英文摘要

Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-supervised rewards derived from its own likelihoods: a direct reward, measuring how likely the constrained output is given the input, and a reverse reward, measuring how well the input can be reconstructed from the generated output. We optimize these rewards with a Reinforce Leave-One-Out (RLOO) objective over groups of grammar-constrained rollouts, augmented with the top-1 beam-search hypothesis and regularized towards a frozen base model. We evaluate GrammarRL on sign language gloss translation, hierarchical text classification, and named entity recognition using Llama models ranging from 1B to 8B parameters. GrammarRL consistently outperforms constrained greedy decoding, with an average improvement of 9.8 points and gains of up to 22.8 BLEU. It matches or outperforms beam search on two of the three tasks while preserving greedy-decoding inference cost. Ablations further show that the two rewards are complementary: either reward alone can underperform the untrained baseline, whereas their combination consistently improves upon it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑