发表机构
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究解决语言模型结构化输出中的空间瓶颈,提出基于自动机路径唯一性的低内存推理引擎,在保持答案一致下实现1.2-2.5倍加速,并大幅降低内存占用。
AI 中文摘要
语言模型越来越频繁地被要求输出结构化内容:遵循模式的 JSON,或带有类型化参数的工具调用。一个小型机器,即自动机,通过禁止会破坏格式的标记来强制执行格式。我们观察到,这台机器具有一个罕见属性:从其任何状态出发,每个标记都恰好沿着一条路径。在图中,只有少数路径连接任意两点,这是复杂度理论中的经典对象,我们的理论结果解决了关于它们的一个开放问题:人们可以决定这样的图是否连接两个点,同时验证它确实具有少量路径,且只需非常少的内存。精确地说,该问题属于类 ReachUL、LOGDCFL、C=L 和 SC2,并且仅需要 O(log2 n / log log n) 空间,低于 Savitch 定理的经典 O(log2 n)。证明背后的构造成为一个推理引擎:格式强制要求的文本无需运行模型即可写出,掩码在 GPU 上无需任何表格即可重新计算,递归格式使用一个小栈,每个输出在标记限制下保持有效,独立字段被并行解码并验证。在配备 Qwen3.5-2B 和 4B 的 16 GB Apple M2 Pro 上,与 MLX 配合 llguidance(该硬件的标准设置)相比,模式约束提取完成时间快 1.2-1.3 倍,且答案相同;一个语法占用 3 MB 而不是高达 1.5 GB;一台服务器可容纳十六个语法,而表格会耗尽内存;十六个工具调用智能体完成时间快 2.5 倍。
英文摘要
Language models are more and more often asked for structured output: JSON that follows a schema, or a tool call with typed arguments. A small machine, an automaton, enforces the format by forbidding the tokens that would break it. We observe that this machine has a rare property: from any of its states, each token leads along exactly one path. Graphs in which only a few paths join any two points are a classical object of complexity theory, and our theoretical result settles an open question about them: one can decide whether such a graph connects two points while verifying that it really has few paths, with very little memory. Precisely, the problem lies in the classes ReachUL, LOGDCFL, C=L and SC2, and needs only O(log2 n/ log log n) space, below the classical O(log2 n) of Savitch's theorem. The constructions behind the proofs become an inference engine: text the format forces is written without running the model, the mask is recomputed on the GPU without any table, recursive formats use a small stack, every output stays valid under a token limit, and independent fields are decoded in parallel and verified. On one 16 GB Apple M2 Pro with Qwen3.5-2B and 4B, against MLX with llguidance, the standard setup for this hardware, schema-constrained extraction finishes 1.2- 1.3x sooner with the same answers, a grammar costs 3 MB instead of up to 1.5 GB, one server holds sixteen grammars where tables run out of memory, and sixteen tool-calling agents finish 2.5x sooner.