arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30922cs.AI

CARVE:扩散语言模型中变长生成的验证式扩展

CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models

Wail Bouhedja, Amr Mohamed, Guokan Shang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出无需重训的 CARVE 算法,解决扩散语言模型固定长度生成的缺陷,通过验证式扩展实现变长生成,在代码生成与数学推理基准中提升性能并降低推理成本。

中文摘要 AI 辅助

掩码扩散语言模型从部分观测的响应画布预测 token,支持双向条件设置与并行 token 细化。然而标准掩码扩散解码器采用固定推理接口:生成开始前已分配给答案的掩码位置数量固定,选择该长度十分困难。较短的画布会截断推理或代码,较长的画布则会浪费计算资源并可能干扰去噪过程。我们提出 CARVE(Counterfactual-Aware Reveal with Verified Expansion,即反事实感知验证式扩展),一种适用于掩码扩散语言模型的无需重新训练的变长算法。CARVE 从较短画布开始,可在解码过程中通过插入额外的 [MASK] 位置来扩展响应。CARVE 并非保留所有插入内容,而是测试候选扩展画布并提出反事实问题:若存在额外的掩码空间,模型会对原始画布中未解决位置做出相似预测吗?仅当插入的掩码在对齐的未解决位置上诱导的 Jensen-Shannon(JS)散度较低时,才保留这些插入的掩码。这使得长度增长成为经过验证的稳定性决策,而非纯粹的置信度启发式方法。CARVE 无需重新训练即可应用于全画布和分块扩散解码器。在代码生成与数学推理基准测试中,CARVE 在所有评估的模型族上始终优于固定长度基线的平均性能。关键的是,CARVE 在实现这些准确率提升的同时降低了推理成本,在部分设置中达到了固定长度解码一半的 FLOPs。

英文摘要

Masked diffusion language models predict tokens from a partially observed response canvas, enabling bidirectional conditioning and parallel token refinement. Yet standard masked-diffusion decoders use a rigid inference interface: the number of masked positions allocated to the answer is fixed before generation begins. Choosing this length is difficult. A short canvas can truncate reasoning or code, while a long canvas wastes computation and can perturb denoising. We introduce CARVE (Counterfactual-Aware Reveal with Verified Expansion), a training-free variable-length algorithm for masked diffusion LMs. Starting from a shorter canvas, CARVE can grow the response during decoding by inserting additional [MASK] positions. Rather than keeping every insertion, CARVE tests a candidate expanded canvas and asks a counterfactual question: would the model make similar predictions for the unresolved positions in the original canvas if the extra masked space were present? The inserted masks are kept only when they induce low Jensen-Shannon (JS) divergence on aligned unresolved positions. This makes length growth a verified stability decision rather than a pure confidence heuristic. CARVE applies without retraining to both full-canvas and blockwise diffusion decoders. Across code generation and mathematical reasoning benchmarks, CARVE consistently improves average performance over fixed-length baselines across all evaluated model families. Crucially, CARVE achieves these accuracy gains while reducing inference cost, reaching half the FLOPs of fixed-length decoding in some settings.

发表机构

  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • Sorbonne Université(索邦大学)
  • Ecole Polytechnique(巴黎综合理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑