arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00605cs.AI

逃离置信陷阱:扩散大语言模型数学推理的进化解码

Escaping Confidence Trap: Evolutionary Decoding for Mathematical Reasoning in Diffusion LLMs

Zhenhong Sun, Hanqing Zhao, Yatao Bian, Rongcheng Tu, Liuyue Xie, Xu Zhang, Jue Wang, Davide Modolo, Daoyi Dong, Dacheng Tao

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散大语言模型数学推理的置信陷阱,提出无训练的进化解码框架,通过逐步选择与分块变异提升LLaDA 2.0的数学推理可靠性。

中文摘要 AI 辅助

扩散大语言模型(dLLMs)作为自回归大语言模型的有前景替代方案,通过分块渐进式去掩蔽实现高效生成,但其通用性能未必能可靠迁移至数学推理领域,数学推理的正确性依赖于连贯的数值-符号推理轨迹。本研究分析LLaDA 2.0的解码轨迹,发现一种反复出现的扩散置信陷阱:在渐进式分块解码过程中,局部 token 的置信度可能与全局推理正确性错位。我们识别出两类典型失败模式:采样敏感型失败,即存在正确路径但不稳定;采样一致型失败,即重复采样收敛至重复的高置信度但错误的延续。基于此,我们提出进化解码(Evolutionary Decoding),这是一种无训练的测试时扩展框架,将扩散解码视为候选推理状态上的进化过程。该框架结合逐步选择(保留有用的数值-符号信号并抑制重复模式)与分块变异(引入结构化替代方案以逃离错误的高置信度区域)。在多个基准上的实验表明,进化解码相较于基于置信度的解码提升了LLaDA 2.0的性能,实现了更可靠的数学推理。

英文摘要

Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive LLMs, offering efficient generation through block-wise progressive unmasking. However, their strong general-purpose performance does not necessarily translate into reliable mathematical reasoning, where correctness depends on preserving coherent numerical-symbolic reasoning trajectories. In this work, we analyze the decoding trajectories of LLaDA 2.0 and identify a recurring diffusion confidence trap: local token confidence can become misaligned with global reasoning correctness during progressive block decoding. Our analysis reveals two representative failure regimes: sampling-sensitive failures, where correct paths exist but are unstable, and sampling-consistent failures, where repeated sampling converges to repetitive high-confidence but incorrect continuations. Motivated by this observation, we propose Evolutionary Decoding, a training-free test-time scaling framework that views diffusion decoding as an evolutionary process over candidate reasoning states. The framework combines step-wise selection, which preserves useful numerical-symbolic signals and suppresses repetitive patterns, with block-wise mutation, which introduces structured alternatives to escape incorrect high-confidence basins. Experiments on multiple benchmarks show that Evolutionary Decoding improves LLaDA 2.0 over confidence-based decoding, leading to more reliable mathematical reasoning.

发表机构

  • Australian National University(澳大利亚国立大学)
  • Nanyang Technological University(南洋理工大学)
  • National University of Singapore(新加坡国立大学)
  • Amazon(亚马逊)
  • University of Technology Sydney(悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑