arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无头部多思维模型

A Model with No Head and Many Thoughts

Nikita Koriagin, Yaroslav Aksenov, George Bredis, Gleb Gerasimov, Nikita Balagansky, Daniil Gavrilov

arXiv 2608.31069首次发表:更新:

发表机构

Yandex; T-Tech(Yandex; T-Tech)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Soft Latent Thinking方法,用轻量投影器替换语言模型头部,在嵌入空间以连续推理步骤替代离散标记,在DeepSeek-Qwen-1.5B和LLaMA-3.2-3B实验中提升了pass@k并减少计算。

AI 中文摘要

大型语言模型在每一步都通过大型词汇头部投射隐藏状态进行解码,该操作计算成本高,且迫使所有推理以离散标记形式表达。我们提出Soft Latent Thinking(软潜在思维)方法,在推理期间将语言模型头部替换为轻量投影器,支持在嵌入空间中进行自回归展开,使推理步骤保持连续而非标记化。在DeepSeek-Qwen-1.5B和LLaMA-3.2-3B上的实验表明,Soft Latent Thinking在所有k值下均持续提升pass@k,同时减少思维链期间的单步计算,在所有软思维方法中达到最高pass@32,证明无需离散标记生成即可在连续空间中进行有效推理。

英文摘要

Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized. Experiments on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B show that Soft Latent Thinking consistently improves pass@k across all k while reducing per-step compute during chain-of-thought. Our method achieves the highest pass@32 among all soft-thinking approaches, demonstrating that effective reasoning can be carried out in continuous space without discrete token generation.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑