发表机构
Qualcomm AI Research; MIT; University of Cambridge; Harvard University(高通人工智能研究部; 麻省理工学院; 剑桥大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LLoCoT框架,以循环Transformer实现非自回归潜在推理,在HumanEval、MBPP上性能优异,较Reasoning SFT大幅缩短延迟并提升吞吐量。
AI 中文摘要
思维链(CoT)推理通过让语言模型在回答前进行额外计算,常可提升其性能。然而,显式CoT将该计算过程表达为自回归生成的 token 序列;潜在推理则用紧凑的连续状态替代这些 token,但多数自回归潜在推理方法仍在潜在向量间保留从左到右的依赖关系。本文提出LLoCoT:一种循环潜在推理框架,它用紧凑潜在工作空间的迭代优化替代从左到右的潜在生成。共享Transformer会被重新应用少量优化迭代,基于提示和不断演化的工作空间状态联合更新潜在槽位。利用优化后的状态,一个概率头预测分布,从中并行采样潜在 token,用于条件化自回归解码器以生成答案。训练采用从显式CoT导出的连续表示,结合最终答案预测损失和基于似然的潜在状态监督。在HumanEval和MBPP数据集上,LLoCoT在评估方法中取得最高均值,准确率与显式CoT基线Reasoning SFT相当,同时优于基础模型、仅答案SFT和NF-CoT;相较于Reasoning SFT,LLoCoT将首个答案token的生成时间缩短约36倍,推理阶段延迟降低约42倍,端到端吞吐量提升9.2%。该设计用并行潜在槽位优化替代串行思维生成,同时保留概率潜在建模和自回归答案解码。
英文摘要
Chain-of-thought (CoT) reasoning often improves language-model performance by giving models additional computation before answering. However, explicit CoT expresses this computation as a sequence of autoregressively generated tokens. Latent reasoning replaces these tokens with compact continuous states, but most autoregressive latent-reasoning methods retain a left-to-right dependency among latent vectors. We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace. A shared transformer is reapplied for a small number of refinement iterations, jointly updating the latent slots based on the prompt and the evolving workspace state. Using the refined state, a probabilistic head predicts a distribution from which latent tokens are sampled in parallel and used to condition an autoregressive decoder for answer generation. Training uses continuous representations derived from explicit CoT together with a final-answer prediction loss and likelihood-based supervision of the latent states. Across HumanEval and MBPP, LLoCoT achieves the highest mean among the evaluated methods, performing on par in accuracy with Reasoning SFT, our explicit-CoT baseline, while outperforming the base model, answer-only SFT and NF-CoT. Relative to Reasoning SFT, LLoCoT reduces time to the first answer token by approximately $36\times$ and reasoning-phase latency by approximately $42\times$, while increasing end-to-end throughput by $9.2\%$. This design replaces serial thought generation with parallel latent-slot refinement while retaining probabilistic latent modeling and autoregressive answer decoding.
Comments9 pages, 1 figure