ELF-REG:将连续扩散语言模型扩展到推理任务
ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks
查看机构详情
- Duke University(杜克大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出ELF-REG,通过表示对齐与纠缠(REPA+REG)扩展连续扩散语言模型至推理任务,在GSM8K、MATH-500、HumanEval和MBPP上超越同规模扩散模型,并支持低NFE下的早停解码。
中文摘要 AI 辅助
完全连续的扩散语言模型(dLMs)在不进行中间离散化的情况下对连续表示进行去噪,然后在最后一步并行解码所有响应令牌。它们在具有挑战性的推理任务上的表现,相较于自回归(AR)大语言模型(LLMs)和掩码扩散语言模型,仍不太成熟。我们将嵌入式语言流(ELF)扩展到GSM8K、MATH-500、HumanEval和MBPP上的数学推理和代码生成任务。我们引入了ELF-REG,它通过表示对齐和纠缠(REPA+REG)来改进学习,其中冻结的自回归教师模型监督中间去噪器特征,并提供与响应联合去噪的全局表示。ELF-REG-L在GSM8K上以64次网络函数评估(NFE)达到55.96%的pass@1,在MATH-500上达到13.39%,在HumanEval上以128 NFE达到22.56%。它在GSM8K和代码的pass@1上优于所评估的规模相当的扩散语言模型,并将MATH-500的pass@1从ELF-L基线的10.55%提高到ELF-REG-L的13.39%。无需少步训练,相同的任务特定检查点通过早停支持强大的低NFE性能,早停在不完成去噪轨迹的情况下解码中间干净预测。在16 NFE下,ELF-REG-L在HumanEval上达到41.21%的pass@10,优于近期规模相当的连续扩散语言模型。
英文摘要
Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and masked dLMs. We scale Embedded Language Flows (ELF) to mathematical reasoning and code generation on GSM8K, MATH-500, HumanEval, and MBPP. We introduce ELF-REG, which improves learning with representation alignment and entanglement (REPA+REG), where a frozen AR teacher supervises intermediate denoiser features and supplies a global representation that is jointly denoised with the response. ELF-REG-L achieves 55.96% pass@1 on GSM8K at 64 network function evaluations (NFE), and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE. It outperforms the evaluated comparable-scale dLMs in pass@1 on GSM8K and code, and improves MATH-500 pass@1 from 10.55% for the ELF-L baseline to 13.39% with ELF-REG-L. Without few-step training, the same task-specific checkpoints support strong low-NFE performance through early-stop, which decodes an intermediate clean prediction without completing the denoising trajectory. At 16 NFE, ELF-REG-L reaches 41.21% HumanEval pass@10, outperforming recent continuous dLMs of comparable scale.