arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35694cs.AI

基于连续潜在扩散的推理

Reasoning with Continuous Latent Diffusion

Xiang Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出潜在流推理模型(LFRMs),通过连续潜在扩散和分阶段课程学习紧凑提示编码器,在数学推理和代码生成任务上超越现有连续扩散基线,并在GSM8K、MATH500、HumanEval等基准上取得显著性能。

中文摘要 AI 辅助

连续扩散通过在潜在空间中进行迭代精炼来生成完整的推理解决方案。我们引入了潜在流推理模型(LFRMs),这是一种基于ELF的训练和推理方案。我们的实验表明,仅靠准确的解码并不能确保强大的推理性能。因此,我们从强自回归教师模型的多个层中学习紧凑的表示。它们的分解还支持以不同速率进行异步去噪。我们表明,提示编码只需保留正确文本条件得分所需的信息,而不必精确匹配教师特征,并使用分阶段课程学习来学习一个紧凑的提示编码器,在推理时替代教师Transformer。我们将DiffusionNFT适应于学习到的自条件引导,并纳入黄金解端点以补充稀疏奖励。我们的监督模型在数学推理和HumanEval代码生成上,在可比骨干规模下,优于近期连续扩散基线的报告结果。使用638M参数的去噪骨干和学习的提示条件,后NFT的LFRM-L在GSM8K上达到63.74%的pass@1,在MATH500上达到24.6%(64步去噪),在HumanEval上达到32.85%,在HumanEval+上达到30.18%(128步去噪)。代码将在以下网址提供:this https URL

英文摘要

Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce the Continuous Embedding Diffusion Reasoner (CEDR), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT CEDR-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/CEDR.

发表机构

  • Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

↑