当潜在推理中的引导失效:潜在到语言的转换差距
When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap
浏览论文内容
中文总结 AI 辅助
本研究揭示潜在推理中激活引导效果弱于显式CoT的现象,提出“潜在到语言转换差距”假设,并通过实验验证,为未来潜在引导方法设计指明关键方向。
中文摘要 AI 辅助
激活引导已成为在显式思维链(CoT)推理过程中控制语言模型的广泛使用的方法,这推动了其向潜在CoT的扩展。然而,我们发现,即使隐藏表示被移动了相当的量,引导连续思维对后续语言生成的影响也远弱于引导显式CoT。我们首先表明,任务信息在连续思维中仍然可识别。因此,我们假设存在一个“潜在到语言的转换差距”,即潜在空间中的干预效果未能转移到语言生成中。两个进一步的结果支持了这一假设:输出分布在转换边界处发生突变,且任务相关方向在潜在CoT中比在显式CoT中发挥的双向控制作用弱得多。这些发现将转换接口确定为评估和设计未来潜在引导方法的核心目标。
英文摘要
Activation steering has become a widely used approach for controlling language models during explicit chain-of-thought (CoT) reasoning, motivating its extension to latent CoT. However, we find that steering continuous thoughts produces substantially weaker effects on subsequent language generation than steering explicit CoT, even when the hidden representations are moved by comparable amounts. We first show that task information remains identifiable in continuous thoughts. Hence, we hypothesize a \textbf{latent-to-language transition gap}, in which an intervention effect in latent space fails to transfer to language generation. Two further results support this hypothesis: the output distribution changes abruptly at the transition boundary, and task-related directions exert much weaker bidirectional control in latent CoT than in explicit CoT. These findings identify the transition interface as a central target for evaluating and designing future latent-steering methods.
发表机构
- HKUSTGZ(香港科技大学(广州))
- SEU(东南大学)
机构由 AI 辅助整理,请以论文原文为准。