AI 中文总结
该研究揭示离散扩散语言模型中确定性PRM引导因弱信号剪枝和最终评判不佳而逊于简单基线,提出保留正确部分解并交由最终状态验证器决策。
AI 中文摘要
离散扩散语言模型(dLLMs)在每一步都暴露一个去噪解,这使得过程奖励模型(PRM)引导看起来像是在测试时消耗计算量的一种方式。我们表明,一旦去噪、PRM评分和结果奖励模型(ORM)评分在相同的前向传播预算内计费,其确定性形式就输给了一个简单得多的基线。我们的PRM对中间去噪状态进行评分,并基于最终答案的正确性进行训练。在Dream-v0-Instruct-7B上,对于GSM8K问题的每个候选8个,每一步保留PRM分数最高的候选达到65.18%,而独立采样加上为该任务训练的ORM重排器达到75.13%。差距在32个候选时扩大到12.69个百分点(pp),在MATH上为9.85 pp,在MBPP上为12.16 pp。我们将其追溯到两个可分离的失败。首先,引导基于弱信号进行剪枝:在GSM8K上,随着掩码比例上升,PRM的ROC-AUC从0.77下降到0.54,当状态用新的rollout重新标记时,这种衰减仍然存在,并且剪枝将候选池中可达到的最佳准确率从独立样本的81.05%降低到67.30%。其次,在GSM8K和MATH上,PRM是一个糟糕的最终评判者:在相同预算下,顺序蒙特卡洛采样器将该上限恢复到77.89%,而使用PRM选择则给出65.48%,与确定性引导相当,而在相同候选上,在最终状态上重新训练的PRM与ORM匹配。MBPP将两者分开:在那里,PRM在重排完成的程序时达到65.47%,与ORM相当,但在引导去噪时只有50.88%。结果指出了dLLM引导的两个目标:在早期去噪过程中保持正确的部分解存活,并将最终选择留给在最终状态上训练的验证器。我们发布了带有结果标签的去噪状态语料库和评估工具包,以便在匹配的计算下进行可复现的比较。
英文摘要
Discrete diffusion language models (dLLMs) expose a denoised solution at every step, which makes process reward model (PRM) guidance look like a way to spend compute at test time. We show that once denoising, PRM scoring, and outcome reward model (ORM) scoring are charged in the same budget of forward passes, its deterministic form loses to a much simpler baseline. Our PRMs score intermediate denoising states and are trained on the correctness of the final answer. On Dream-v0-Instruct-7B with 8 candidates per GSM8K problem, keeping the candidate with the highest PRM score at every scoring step reaches 65.18%, while independent sampling plus an ORM reranker trained for the task reaches 75.13%. The gap grows to 12.69 percentage points (pp) with 32 candidates, and is 9.85 pp on MATH and 12.16 pp on MBPP. We trace it to two separable failures. First, guidance prunes on a weak signal: on GSM8K, PRM ROC-AUC falls from 0.77 to 0.54 as the mask ratio rises, a decay that persists when states are relabeled with fresh rollouts, and pruning lowers the best accuracy reachable from the candidate pool from 81.05% for independent samples to 67.30%. Second, on GSM8K and MATH, the PRM is a poor final judge: a sequential Monte Carlo sampler at the same budget restores that ceiling to 77.89%, yet selecting with the PRM gives 65.48%, on par with deterministic guidance, while a PRM retrained on final states matches the ORM on identical candidates. MBPP separates the two: there the PRM reaches 65.47% when reranking finished programs, on par with the ORM, but 50.88% when it guides denoising. The results point to two targets for dLLM guidance: keep correct partial solutions alive through early denoising, and leave the final choice to a verifier trained on final states. We release the corpus of denoising states with outcome labels and evaluation toolkit for reproducible comparisons at matched compute.
CommentsAccepted at NeurIPS 2026. 27 pages, 6 figures. Code: https://github.com/dLLM-PRM-Gap/dLLM-PRM-Gap; dataset and model: https://huggingface.co/collections/YanZhanPKU/dllm-prm-gap