发表机构
Max Planck Institute for Intelligent Systems; ETH Zurich; ELLIS Institute Tübingen; University of Oxford; Tübingen AI Center; Liquid AI(马克斯·普朗克智能系统研究所; 苏黎世联邦理工学院; ELLIS图宾根研究所; 牛津大学; 图宾根人工智能中心; Liquid AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究流语言模型(FLMs)的连续表示在推理中的作用,通过理论和实验证明其在少步去噪下比离散扩散模型更高效,且可用更小模型达到同等准确率。
AI 中文摘要
流语言模型(FLMs)已成为离散扩散语言模型的一种连续状态替代方案,但其连续表示在推理中的作用仍不清楚。我们通过比较FLMs与离散扩散模型在匹配去噪步数下的求解准确率来研究这一问题。与离散扩散在去噪步骤之间传递分类状态不同,FLMs在整个去噪过程中演化连续序列表示,并仅在最后将其解码为离散标记。我们的理论分析从叠加视角展示了这些连续状态中保留的信息如何有益于推理。中间状态干预为这一理论解释提供了进一步的实证支持,表明移除关于备选候选的信息会降低后续解的恢复能力。综合这些发现,FLMs允许多个候选的证据持续存在,并在产生离散答案之前为后续推理提供信息。此外,我们在迷宫规划和数独任务上的实验表明,FLMs在少步数场景下实现了更高的推理效率:在匹配模型大小和较小去噪步数下,FLMs的序列准确率高于离散扩散基线。在迷宫规划任务中,FLMs还能以更小的模型达到相当的准确率。例如,在Maze15上,FLM以比MDLM少36.5%的参数,在64步去噪时达到95%的准确率目标。这些发现表明,连续状态空间是推理模型的一个有前景的基础,此类模型需要更少的精炼步骤。
英文摘要
Flow Language Models (FLMs) have emerged as a continuous-state alternative to discrete diffusion language models, yet the role of their continuous representations in reasoning remains unclear. We investigate this question by comparing the reasoning efficiency of FLMs and discrete diffusion models, measured by solution accuracy under matched denoising steps. Unlike discrete diffusion, which passes categorical states between denoising steps, FLMs evolve a continuous sequence representation throughout denoising and decodes it into discrete tokens only at the end. Our theoretical analysis shows, from a superposition perspective, how information retained in these continuous states can benefit reasoning. Intermediate-state interventions provide further empirical support for this theoretical account, showing that removing information about alternative candidates reduces subsequent solution recovery. Together, these findings show that FLMs allow evidence for multiple candidates to persist and inform subsequent reasoning before a discrete answer is produced. Furthermore, our experiments on maze planning and Sudoku tasks show that FLMs achieve greater reasoning efficiency in the few-step regime: FLMs achieves higher sequence accuracy than discrete diffusion baselines at matched model sizes and small denoising steps. On maze planning tasks, FLMs can also achieve comparable accuracy with smaller models. For example, on Maze15, FLM reaches the 95\% accuracy target at 64 denoising steps with 36.5\% fewer parameters than MDLM. These findings point to continuous state spaces as a promising foundation for reasoning models that require fewer refinement steps.