arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度思考,一次输出:ReLIT,一种递归隐式隐式Transformer框架

Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework

Abhishek Panwar, Maheep Singh, Saksham Bansal

arXiv 2608.08113首次发表:更新:

发表机构

Indian Institute of Technology Roorkee(鲁尔基印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对思维链提示计算开销大、微型递归模型语义连贯性差的问题,提出ReLIT框架,结合TinyLlama-1.1B骨干与可训练递归块,在逻辑推理基准上高效实现推理,性能优于或媲美更大模型。

AI 中文摘要

思维链(CoT)提示已成为诱导大型语言模型(LLMs)进行推理的主流范式,但它要求模型将中间推理步骤作为离散标记外化,从而产生大量计算开销。近期的隐式推理方法试图将这一过程内化到连续隐藏状态中。该领域最新进展之一是微型递归模型(TRMs),其擅长符号推理,但在自然语言场景中难以保持语义连贯性。为弥合这一差距,我们提出ReLIT(Recursive Latent Implicit Transformer,递归隐式隐式Transformer),这是一种将深度递归推理建立在基础模型丰富语义表示之上的混合框架。ReLIT在冻结的LLM骨干(TinyLlama-1.1B)基础上,增加了一个轻量级可训练递归块,该块会迭代优化其隐式思考(z),然后才生成最终输出,从结构上解决了算法处理中的语言直觉问题,并通过梯度隔离的循环回路实现“深度思考”,且无显式标记生成的延迟。实验表明,ReLIT在GLoRE逻辑推理基准上实现了高参数效率,在ProofWriter和RuleTaker等具有挑战性的任务上,尽管仅使用最少监督,其性能可与大得多的模型相匹配甚至超越。这些结果表明,推理能力可通过递归深度而非参数宽度高效扩展,为基于语义的隐式推理提供了一个原则性框架。

英文摘要

Chain-of-Thought (CoT) prompting has become the dominant paradigm for eliciting reasoning in Large Language Models (LLMs), yet it creates substantial computational overhead by forcing models to externalize intermediate reasoning steps as discrete tokens. Recent latent reasoning approaches attempt to internalize this process within continuous hidden states. One of the latest advancements in the field of latent reasoning, Tiny Recursive Models (TRMs) excel at symbolic reasoning but struggle to preserve semantic coherence in natural language settings. To bridge this gap, we introduce ReLIT (Recursive Latent Implicit Transformer), a hybrid framework that grounds deep recursive reasoning within the rich semantic representations of a foundational model. ReLIT augments a frozen LLM backbone (TinyLlama-1.1B) with a lightweight, trainable recursive block that iteratively refines its latent thinking (z) before committing to a final output, structurally solving linguistic intuition from algorithmic processing and enabling "deep thinking" via gradient-isolated recurrent loops without the latency of explicit token generation. Empirically, ReLIT achieves high parameter efficiency on the GLoRE logical reasoning benchmark, matching or outperforming significantly larger models on challenging tasks such as ProofWriter and RuleTaker despite minimal supervision. These results demonstrate that reasoning capability can be scaled efficiently through recurrent depth rather than parameter width, offering a principled framework for semantically grounded implicit reasoning.

Comments14 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑