arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28272cs.CL

迈向高效推理:为扩散语言模型学习因果捷径

Towards Efficient Reasoning: Learning Causal Shortcuts for Diffusion Language Models

  • Zhejiang University(浙江大学)
  • Shanghai AI Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Dian Jin, Kairong Han, Baohong Li, Xinpeng Dong, Zijing Hu, Nuanqiao Shan, Fei Wu, Kun Kuang

AI总结:

针对扩散语言模型在双向注意力下推理效率低的问题,提出因果捷径学习框架,通过提取并优先掩码引导令牌,显著提升推理准确性和收敛速度,在多个基准上超越现有基线。

AI中文摘要:

扩散语言模型(DLMs)因其强大的推理能力而受到广泛关注。然而,在双向注意力机制下,与自回归模型(ARMs)相比,DLMs 在指数级增大的探索空间中运行,这使得在随机掩码下难以聚焦于引导推理的令牌。我们将因果捷径定义为覆盖完整序列并为正确推理轨迹提供明确引导的令牌链。我们分析了因果捷径对 DLMs 推理准确性和收敛速度的影响,发现它们大幅提高了答案收敛效率和生成准确性。受此启发,我们提出了一个用于 DLMs 的因果捷径学习(CSL)框架。具体来说,我们引入了一个逐步令牌提取过程,从数据中提取因果捷径,并在训练期间对这些令牌应用并行优先级掩码,从而通过因果捷径实现高效且准确地向正确答案收敛。在多个推理基准和两个基础模型上的大量实验表明,CSL 持续优于现有的 SFT 变体基线,在仅 SFT 的模型上平均提高了 1.92%,在 MATH-500 上最高提高了 4.20%。代码可在 https://github.com/ZJUDianJin/Causal-Shortcuts-Learning 获取。

英文摘要:

Diffusion Language Models (DLMs) have attracted significant attention for their strong reasoning ability. However, under a bidirectional attention mechanism, DLMs operate over an exponentially large exploration space compared to autoregressive models (ARMs), making it challenging to focus on reasoning-guiding tokens under random masking. We define causal shortcuts as token chains that cover the full sequence and provide explicit guidance towards correct reasoning trajectories. We analyze the effects of causal shortcuts on the reasoning accuracy and convergence speed of DLMs, and find that they largely improve answer convergence efficiency and generation accuracy. Motivated by this, we propose a Causal Shortcut Learning (CSL) Framework for DLMs. Specifically, we introduce a step-by-step token extraction procedure to extract causal shortcuts from data, and apply parallel prioritized masking on these tokens during training to enable efficient and accurate convergence to correct answers via causal shortcuts. Extensive experiments across multiple reasoning benchmarks and two base models demonstrate that CSL consistently outperforms existing SFT-variant baselines, achieving an average improvement of $1.92\%$ over SFT-only models, and up to $4.20\%$ on MATH-500. The code is available at the \href{https://github.com/ZJUDianJin/Causal-Shortcuts-Learning}{https://github.com/ZJUDianJin/Causal-Shortcuts-Learning

↑