发表机构
Yandex; T-Tech(Yandex; T-Tech)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Soft Latent Thinking方法,用轻量投影器替换语言模型头部,在嵌入空间以连续推理步骤替代离散标记,在DeepSeek-Qwen-1.5B和LLaMA-3.2-3B实验中提升了pass@k并减少计算。
AI 中文摘要
大型语言模型在每一步都通过大型词汇头部投射隐藏状态进行解码,该操作计算成本高,且迫使所有推理以离散标记形式表达。我们提出Soft Latent Thinking(软潜在思维)方法,在推理期间将语言模型头部替换为轻量投影器,支持在嵌入空间中进行自回归展开,使推理步骤保持连续而非标记化。在DeepSeek-Qwen-1.5B和LLaMA-3.2-3B上的实验表明,Soft Latent Thinking在所有k值下均持续提升pass@k,同时减少思维链期间的单步计算,在所有软思维方法中达到最高pass@32,证明无需离散标记生成即可在连续空间中进行有效推理。
英文摘要
Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized. Experiments on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B show that Soft Latent Thinking consistently improves pass@k across all k while reducing per-step compute during chain-of-thought. Our method achieves the highest pass@32 among all soft-thinking approaches, demonstrating that effective reasoning can be carried out in continuous space without discrete token generation.
CommentsAccepted to Findings of EMNLP 2026