在潜空间思考,以语言解释:自解释潜推理
Think in Latent, Explain in Language: Self-Explainable Latent Reasoning
AI总结:
本研究提出自解释潜推理框架SELR,通过多任务训练使单个模型兼具高效潜推理与自解释能力,在LLMs和VLMs上实现更优效率、准确性与自包含可解释性。
AI中文摘要:
潜推理已成为基于文本的思维链(Chain-of-Thought, CoT)的有力替代方案,它通过将冗长的推理压缩为紧凑的嵌入表示,在计算效率上取得了显著提升。然而,将推理压缩到潜空间会使思考过程变得不透明,阻碍了其可解释性。当前方法存在明显的权衡:它们要么是无法解释的“黑箱”(例如Coconut),其潜推理过程不具备人类可读性;要么依赖单独的事后解码器实现可解释性(例如Heima),这会引入架构开销,并使解释与实际推理过程解耦。在本研究中,我们提出了一个用于自解释潜推理(Self-Explainable Latent Reasoning, SELR)的统一框架,该框架训练单个模型以执行高效且固有可解释的潜推理。我们的核心贡献是一种新颖的多任务训练目标,同时优化两个目标:(1)答案损失,用于优化潜推理轨迹以生成准确的最终答案;(2)CoT损失,用于显式训练同一模型将其自身的潜表示解码为人类可理解的推理步骤。该设计确保生成的潜表示既具有任务有效性,又具有语义可解释性,无需外部解码器。我们在大语言模型(Large Language Models, LLMs)和视觉语言模型(Vision-Language Models, VLMs)上验证了SELR的有效性,结果表明,与基线方法相比,SELR实现了更优的令牌效率和准确性,同时独特地提供了无需辅助模型的自包含可解释性。项目页面可访问此URL。
英文摘要:
Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing reasoning into the latent space renders the thinking opaque, hindering its interpretability. Current methods present a stark trade-off: they either function as unexplainable ''black boxes'' (e.g., Coconut), where the latent reasoning is not human-readable, or rely on separate post-hoc decoders for explainability (e.g., Heima), introducing architectural overhead and decoupling the explanation from the actual reasoning process. In this work, we present a unified framework for Self-Explainable Latent Reasoning (SELR) that trains a single model to perform efficient and inherently explainable latent reasoning. Our core contribution is a novel multi-task training objective that optimizes for two goals simultaneously: (1) an Answer Loss that optimizes the latent reasoning trajectory to produce accurate final answers, and (2) a CoT Loss that explicitly trains the same model to decode its own latent representations back into human-understandable reasoning steps. This design ensures that generated latent representations are both task-effective and semantically interpretable, eliminating the need for external decoders. We validate the effectiveness of SELR on both Large Language Models (LLMs) and Vision-Language Models (VLMs), demonstrating that SELR achieves superior token efficiency and accuracy compared to baselines, while uniquely providing self-contained explainability without auxiliary models. Project page is available at https://jasondayuan.github.io/SELR/.