通过潜在空间中的约束优化实现高效推理
Efficient Reasoning via Constrained Optimization in Latent Space
浏览论文内容
中文总结 AI 辅助
针对大型推理模型过度思考导致的令牌浪费,提出基于潜在空间约束优化的免训练框架,通过二次规划约束推理步骤,在四个模型和六个基准上实现最高12.1%的准确率提升并减少11.8%-52.8%的令牌生成。
中文摘要 AI 辅助
大型推理模型(LRMs)展现出卓越的推理能力,但仍存在过度思考的问题,即生成冗余的推理步骤,导致大量令牌消耗。现有方法,如抑制反思性关键词或强制缩短推理长度,试图缓解这一问题,但不可避免地截断必要步骤并引发思考不足,从而损害性能。为解决这一困境,我们研究了潜在表示,并观察到高效推理步骤在潜在空间中自然聚集于一个集中区域,而偏离该区域的步骤往往产生冗长的序列。为利用这一点,我们通过一个二次规划将偏离的隐藏状态投影回该区域,从而使推理保持在该区域内聚焦。随后,我们提出了一种新颖的免训练框架,以实现高效推理,在不牺牲性能的情况下降低令牌生成成本。在从1.5B到14B的四个模型以及数学推理、编码和科学问答的六个基准上进行的广泛实验验证了我们方法的有效性,准确率最高提升12.1%,同时生成的令牌减少11.8%至52.8%。代码可在以下网址获取:this https URL。
英文摘要
Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1\% improvement in accuracy while reducing generated tokens by 11.8\% to 52.8\%. Codes are available at \href{https://github.com/hzn18/Opt4Reasoning}{https://github.com/hzn18/Opt4Reasoning}.