arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11951cs.AIcs.CLcs.FLcs.PL

GRID:用于企业SQL生成的语法约束解码

GRID: Grammar-Railed Decoding for Enterprise SQL Generation

Mohsen Arjmandi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对企业SQL生成需求,提出GRID语法约束解码引擎,通过特定方式确定令牌掩码,编译角色访问控制,有四个保证。经实验,Rust内核降低掩码时间,在Spider上约束解码提升模型可执行率,还具备审计跟踪和篡改检测功能。

中文摘要 AI 辅助

大语言模型可以编写SQL,但企业部署需要的不仅仅是合理的文本:输出必须在语法上有效,必须符合每个角色和每个模式的策略,必须提供可证明的(而非尽力而为的)保证,不能随着生成数量的增加而变慢,并且必须为每个决策留下合规级别的记录。我们提出了GRID(语法约束解码),这是一种语法约束解码引擎,它根据解析器配置(词法分析器扫描状态x LALR(1)栈)而不是令牌序列来确定确切的下一个令牌掩码,并使用逐步推进的LALR(1)解析器本身作为可行前缀预言机。大语言模型令牌通过字节级前缀树遍历与语法终端桥接,通过独立于上下文/依赖于上下文的分割,使得缓存键的健全性在构造时成立。基于角色的访问控制被编译到语言中:角色投影使语法的产生式子集化,模式词典限制标识符终端,因此在掩码级别无法访问被禁止的动词和标识符。阐述了四个保证(健全性、完整性、终止性和近乎恒定的每个令牌成本),并带有明确的前提条件,每个条件都与一个测试或基准配对。Rust内核将每个令牌掩码的中位数降低到3.6 - 6.7微秒,在两个令牌化器上的p50和p90处领先于llguidance,且零误判。在Spider上,约束解码在0.5B时价值+13个执行准确率点,并且对可证明的掩码不可强制执行的残余物(列级策略)进行一次检查器引导的修复过程,可将7B模型的可执行率提高到94.5%。一个哈希链每个令牌审计跟踪以相同的位进行重放,并具有100%的篡改检测。我们明确说明了掩码不能做的事情(分布忠实性、列级RBAC、非LALR(1)语言)以及测量成本仍然存在的地方。

英文摘要

Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-role and per-schema policy, must carry provable (not best-effort) guarantees, must not slow down as generations grow, and must leave a compliance-grade record of every decision. We present GRID (Grammar-Railed Decoding), a grammar-constrained decoding engine that keys exact next-token masks on parser configurations (lexer scan state x LALR(1) stack) rather than on token sequences, and uses the incrementally advanced LALR(1) parser itself as a viable-prefix oracle. LLM tokens are bridged to grammar terminals by a byte-level trie walk with a context-independent/context-dependent split that makes cache-key soundness hold by construction. Role-based access control is compiled into the language: role projections subset the grammar's productions and schema lexicons restrict identifier terminals, so forbidden verbs and identifiers are unreachable at mask level. Four guarantees (soundness, completeness, termination, and near-constant per-token cost) are stated with explicit preconditions and each paired with a test or benchmark. Rust kernels bring the per-token mask to a 3.6-6.7 us median, ahead of llguidance at p50 and p90 on two tokenizers with zero false rejects; per-token guard cost is position-flat at n=16,000. On Spider, constrained decoding is worth +13 execution-accuracy points at 0.5B, and one checker-guided repair pass over the provably mask-unenforceable residue (column-level policy) lifts a 7B model to 94.5% executable. A hash-chained per-token audit trail replays bit-identically with 100% tamper detection. We state plainly what the mask cannot do (distribution faithfulness, column-level RBAC, non-LALR(1) languages) and where measured cost remains.

发表机构

  • evolutionID GmbH(进化ID有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑