arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37976cs.LGcs.AIcs.CLcs.CV

$S^3$:谱零空间交换使推理模型高效

$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient

发表机构伊利诺伊大学厄巴纳-香槟分校 · 清华大学 · 达特茅斯学院
另 1 家 · 查看机构详情
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Tsinghua University(清华大学)
  • Dartmouth College(达特茅斯学院)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang, Ge Liu

首次发表
浏览论文内容

中文总结 AI 辅助

提出谱零空间交换($S^3$),利用思考与非思考模型权重在零空间的分量差异,无需训练即可在保持准确率的同时平均减少27.4%推理令牌开销,并提升1.0%任务准确率。

中文摘要 AI 辅助

使用思维链训练的LLM在推理能力上表现出色,但往往伴随过高的令牌成本。我们发现推理能力的核心在于思考模型权重分量中,由对应非思考模型的主导奇异方向所定义的投影的零空间内,而移除该子空间分量可以大幅提升推理效率,同时不损害思维模式后训练期间获得的准确性。与现有主要作用于主导子空间的工作不同,我们首次揭示了零空间的关键作用并将其用于模型优化。基于这一发现,我们提出谱零空间交换($S^3$),这是一种无需训练的对配对非思考与思考检查点的组合方法。我们的方法将非思考模型保持在其自身的主导子空间内,并将思考检查点置于该子空间之外,从而在保持准确性的同时提升推理效率。我们在2B-30B的稠密和混合专家(MoE)架构上广泛评估了$S^3$,覆盖数学、多模态和音频推理领域的28个评估环境。$S^3$在无训练模型组合策略中建立了新的经验帕累托前沿:在所有设置中,与完整思考模型相比,它平均减少推理令牌开销27.4%,同时整体任务准确率提高1.0个百分点(例如,在HMMT25上获得+8.3%的准确率提升,同时实现33.0%的令牌加速)。我们进一步使用注意力熵进行解释,发现保留的分量产生更集中的注意力,并使用一个简化的优化分析模型来证明零空间为何能有效降低注意力熵,从而提升推理效率。

英文摘要

LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.

补充信息

↑