arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DCGC:基于草稿条件的全局修正,用于带掩码扩散模型的复杂推理

DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models

Minhae Oh, Nakyung Lee, Jungwoo Lee

arXiv 2608.25428首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DCGC是结合SFT与动态双CFG机制的MDM框架,用于修正LLMs推理错误,在多推理基准中优于现有方法,可作为无验证器的全局修正模块。

AI 中文摘要

修正有缺陷的推理轨迹仍是大型语言模型(LLMs)面临的重大挑战,其自回归生成会将早期错误传播到后续推理中。我们引入DCGC,这是一种掩码扩散模型(MDM)框架,用于全局修正,它将上游求解器生成的不完善解决方案草稿作为辅助上下文。DCGC结合了特定任务的监督微调(SFT)与一种名为动态双分类器引导(Dynamic Dual-CFG)的新型推理时机制,该机制将仅问题分支与问题-草稿联合分支分离,并使用相对置信度差距缩放草稿条件残差。在数学、代码和知识推理基准测试中,DCGC的性能优于标准采样和更简单的CFG变体;额外结果表明其可迁移至不同的扩散骨干网络。在测试时无真实失败标签的设置下,DCGC通过修正低共识的上游输出提升了整个测试集的准确率,凸显了其作为无验证器的全局修正模块在困难推理实例中的实用性。

英文摘要

Correcting flawed reasoning traces remains a significant challenge for Large Language Models (LLMs), whose autoregressive generation can propagate early mistakes into subsequent reasoning. We introduce DCGC, a Masked Diffusion Model (MDM) framework for global correction that uses an imperfect solution draft from an upstream solver as auxiliary context. DCGC combines task-specific Supervised Fine-Tuning (SFT) with a novel inference-time mechanism called Dynamic Dual-CFG. This mechanism separates problem-only and joint problem-draft branches and scales the draft-conditioned residual using a relative confidence gap. Across math, code, and knowledge reasoning benchmarks, DCGC outperforms standard sampling and simpler CFG variants, with additional results suggesting transfer to different diffusion backbones. In test-time setting where ground-truth failure labels are unavailable, DCGC improves full test set accuracy by correcting low-consensus upstream outputs, highlighting its utility as a verifier-free global correction module for difficult reasoning instances.

Comments21 pages, 3 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑