发表机构
University of Oxford; Stanford University(牛津大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对语言模型推理中仅依赖自然语言受限的问题,提出基于J空间的J-CoT循环推理框架,通过词汇索引系数传递中间状态,在多任务中表现出色,J-CoT-Zero匹配或超基线,J-CoT-Train获高分。
AI 中文摘要
思维链提示通过在连续计算步骤中传递中间状态来改进语言模型推理。然而,仅依赖自然语言作为唯一的循环接口过于受限,因为许多瞬时计算无需完全语言化。现有潜在推理方法通过循环传播连续隐藏状态消除此约束,但整体传递密集隐藏向量,缺乏明确机制选择和组织下一个推理步骤所需信息。我们引入J-CoT,一个基于J空间的循环推理框架,J空间是模型隐藏表示内基于词汇索引的坐标系。在每个循环中,模型在其完整隐藏空间中计算。在循环边界,J-CoT将中间状态表示为词汇索引系数,作为J思维向前传递,并映射回模型隐藏表示用于下一个循环。因此,J-CoT既不需要流畅的中间推理依据,也不需要对完整隐藏状态进行循环。在匹配的骨干和推理设置下,J-CoT-Zero在每个基准测试中匹配或超过最强的评估潜在推理基线,而J-CoT-Train在评估的数学、科学、编码和结构化路径推理任务中获得最高分。
英文摘要
Chain-of-thought prompting improves language-model reasoning by carrying intermediate states across successive computation steps. However, relying on natural language as the only recurrent interface is overly restrictive, since many transient computations do not need to be fully verbalized. Existing latent-reasoning methods remove this constraint by recurrently propagating continuous hidden states. However, these methods pass a dense hidden vector as a whole, without an explicit mechanism for selecting and organizing the information needed by the next reasoning step. This motivates an intermediate interface that remains linguistically grounded without requiring a decoded sentence. We introduce \textbf{J-CoT}, a recurrent reasoning framework built on \emph{J-space}, a vocabulary-indexed coordinate system within the model's hidden representations. Within each cycle, the model computes in its full hidden space. At the cycle boundary, J-CoT expresses the intermediate state as vocabulary-indexed coefficients, carries these coefficients forward as a \emph{J-thought}, and maps them back into the model's hidden representation for the next cycle. J-CoT therefore requires neither a fluent intermediate rationale nor recurrence over the complete hidden state. Under matched backbone and inference settings, J-CoT-Zero matches or exceeds the strongest evaluated latent-reasoning baseline on every benchmark, while J-CoT-Train obtains the highest score across the evaluated mathematical, scientific, coding, and structured path-reasoning tasks.
Commentswork in progress