arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

约束是图,而非链:扩散语言模型的精确解码

Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models

Jianchang Su, Wei Zhang

arXiv 2609.32900首次发表:更新:

AI 中文总结

针对扩散语言模型约束解码中顺序编码的指数膨胀问题,提出FactorDLM,以因子图表示关系约束并通过变量消除精确条件化,在九个基准上实现零违规且开销低,并证明编码选择可预先计算。

AI 中文摘要

扩散语言模型(dLLMs)以任意顺序预测掩码位置,但其精确约束解码器仍将约束编码为顺序语言,其状态必须跟踪位置之间每个未解析的依赖关系。对于关系约束,这种编码呈指数级增长:对于同序复制,每个有限自动机(无论是确定性的还是非确定性的)需要$4^k$个状态,每个上下文无关文法的大小为$2^{\Omega(k)}$,而同一关系的因子图大小为$O(k)$,且峰值表仅含16个条目。我们提出FactorDLM,一种免训练解码器,它将有限域关系表示为因子图,并在每个去噪步骤中,通过变量消除将模型的平均场预测精确地条件化于该图。解码成本随约束图的诱导宽度呈指数增长,该宽度取代自动机大小成为主导参数。由于有限自动机是链状因子图,一个编译器即可同时强制执行语法和非局部关系:在具有跨字段引用的JSON记录上,仅使用模式自动机会导致引用悬空,仅使用关系因子会产生格式错误的JSON,而组合方案在两方面均有效,包括变长记录。在九个关系基准和三个骨干网络上,每个输出在0.4%-6.9%的投影开销下满足所有声明的约束,而无约束解码的有效性为0%-79%;编译后的投影在重复查询上比使用八个并行工作线程的CP-SAT快13.6倍。由于无模型规则解决了五个标准基准中的三个,我们构建了具有精确随机基线和固定模板下限的基准,在这些基准上,在精确约束样本中进行选择优于贪心投影。哪种编码更廉价——顺序状态还是直接因子——取决于约束,且可在解码开始前计算。

英文摘要

Diffusion language models (dLLMs) predict masked positions in arbitrary order, but their exact constrained decoders still encode constraints as sequential languages, whose state must track every unresolved dependency between positions. For relational constraints this encoding grows exponentially: for same-order copy, every finite automaton needs $4^k$ states, deterministic or nondeterministic, and every context-free grammar has size $2^{Ω(k)}$, while the factor graph of the same relation has size $O(k)$ and a 16-entry peak table. We introduce FactorDLM, a training-free decoder that represents finite-domain relations as a factor graph and, at each denoising step, conditions the model's mean-field prediction on that graph exactly by variable elimination. Decoding cost then grows exponentially with the induced width of the constraint graph, which replaces automaton size as the governing parameter. Because a finite automaton is a chain-shaped factor graph, one compiler enforces syntax and nonlocal relations together: on JSON records with cross-field references, a schema automaton alone leaves references dangling, relational factors alone produce malformed JSON, and the combined plan is valid on both counts, including on records of variable length. Across nine relational benchmarks and three backbones, every output satisfies every declared constraint at 0.4-6.9% projection overhead, where unconstrained decoding is 0-79% valid, and compiled projection answers repeated queries 13.6x faster than CP-SAT with eight parallel workers. Because model-free rules solve three of five standard benchmarks, we construct benchmarks with exact chance and fixed-template floors, on which selecting among exact constrained samples beats greedy projection. Which encoding is cheaper, sequential state or direct factors, depends on the constraint and is computable before decoding begins.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑