arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过带保障的声明式智能体编程学习用于语法约束解码的上下文无关文法

Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees

Kevin Cheang, Geoff Hulette, Rahul Kumar, Felipe R. Monteiro, Federico Mora, Robin Salkeld, Lin Tan, Serdar Tasiran

arXiv 2608.05493首次发表:更新:

发表机构

Amazon Web Services; Purdue University(亚马逊云科技; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出名为Autogrammar的声明式智能体,可从文档和执行数据自动学习上下文无关文法,用于语法约束解码,在三种DSLs上的实验显示其生成的文法性能优于现有基线,能显著提升端到端LM的真实任务表现。

AI 中文摘要

语言模型(LMs)越来越多地通过用领域特定语言(DSLs)编写的程序与外部服务交互。遗憾的是,由于DSLs通常资源匮乏且鲜为人知,LMs经常生成这些语言中语法无效的程序。语法约束解码可消除此类失败,但需要语法约束,这些约束通常采用目标语言的上下文无关文法形式,而第三方DSLs很难获得这种人工制品。在本研究中,我们定义了一个名为Autogrammar的智能体,它可从文档和执行数据中自动学习上下文无关文法。Autogrammar被形式化为Kripke结构,其非确定性选择由语言模型解决,从而通过线性时序逻辑约束实现对智能体行为的声明式控制。我们在三种DSLs(即Amazon CloudWatch Logs Insights、Dynatrace Query Language和Datadog Search Syntax)上评估了四个版本的Autogrammar,发现它生成的文法在未见数据上达到近乎完美的精度;时序约束将执行时间减少3.8倍,且未造成统计学上显著的精度损失;执行数据至关重要,而文档并非必需;使用Autogrammar生成的文法进行语法约束解码,在10个真实任务中的8个任务上显著提高了端到端LM性能,达到或超过了专业维护文法的性能。相比之下,现有LM基线和最先进形式技术生成的上下文无关文法在相同评估中表现明显更差。

英文摘要

Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately, since DSLs are often low-resource and esoteric, LMs frequently produce syntactically invalid programs in these languages. Grammar-constrained decoding can eliminate such failures, but requires syntactic constraints. These are usually in the form of a context-free grammar for the target language, an artifact that is hard to come by for third-party DSLs. In this work, we define an agent, called Autogrammar, that automatically learns context-free grammars from documentation and execution data. Autogrammar is formalized as a Kripke structure whose nondeterministic choices are resolved by a language model, enabling declarative control of agent behavior via linear temporal logic constraints. We evaluate four versions of Autogrammar on three DSLs (i.e., Amazon CloudWatch Logs Insights, Dynatrace Query Language, and Datadog Search Syntax) and find that it generates grammars that achieve near perfect precision on unseen data; that temporal restrictions reduce execution time by 3.8x without incurring statistically-significant loss in precision; that execution data is crucial while documentation is dispensable; and that grammar-constrained decoding using Autogrammar-generated grammars significantly improves end-to-end LM performance on eight out of ten real tasks, matching or exceeding the performance of a professionally-maintained grammar. In comparison, the context-free grammars generated by existing LM baselines and a state-of-the-art formal technique perform significantly worse over the same evaluation.

Comments9 pages, 3 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑