arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06666cs.LGcs.CL

流匹配潜在推理的关键因素

What Matters for Latent Reasoning with Flow Matching

Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出流式潜在推理(FLaRe),通过流匹配在潜在空间训练有效思维,满足有用、多样、可解释、可改进和高效五项要求,在算术基准上以四分之一延迟达到显式思维链97%的准确率。

中文摘要 AI 辅助

潜在推理让大型语言模型(LLM)在连续空间中思考,并仅将答案用语言表达出来。我们认为,一个有效的潜在思维必须满足五个要求:它应当是有用的,有助于产生正确答案而非仅仅改变答案;应当是多样的,以便重新采样能产生不同的推理轨迹;应当是可解释的,使得解码出的思维链(CoT)能反映答案实际遵循的推理过程;应当是可通过更多推理计算来改进的;并且应当是高效的,在相近准确率下,其成本低于显式思维链。当前方法很少满足这些要求:它们从问题中学习捷径,将显式思维链蒸馏到模型权重中,或逐词模仿它。我们专注于在学得的潜在空间中进行流匹配,我们认为这一族方法最适合满足这些要求,并确定了使其有效的训练选择。其成果是流式潜在推理(FLaRe),一个涵盖潜在空间编码内容及其塑造方式、流训练位置、答案读出方式以及最终基于模型自身验证思维进行训练阶段的简单方案。针对每项要求的探针测试表明,FLaRe在全部五个方面均优于先前的潜在方法。在算术基准上,它也表现更优,同时以显式思维链四分之一的延迟达到了其97%的准确率。

英文摘要

Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.

↑