发表机构
School of Artificial Intelligence, Jilin University; The Hong Kong Polytechnic University(吉林大学人工智能学院; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于稀疏自编码器分析DeepSeek-R1-Distill-Qwen-7B的推理机制,揭示思考与非思考模式的神经差异,为大语言模型推理可解释性提供新见解。
AI 中文摘要
尽管采用思维链(Chain-of-Thought,CoT)的大语言模型(Large Language Models,LLMs)展现出卓越的推理能力,但区分这种显式思考模式与直接生成答案的非思考模式的神经机制仍鲜为人知。为解构这一认知过程,我们将Top-K稀疏自编码器(Sparse Autoencoders,SAEs)应用于DeepSeek-R1-Distill-Qwen-7B的中间表示,并考察模型在三种不同难度的数学求解任务中的不同行为。观察发现,两种推理模式下模型运行方式存在明显差异:思考模式依赖稀疏且高强度的特征激活驱动与问题复杂度无关的文字推导,而非思考模式则呈现自适应且弥散的模式,优先进行符号操作。因果分析显示,通过总激活量抑制三个最活跃的稀疏特征揭示了三项原则:(i)推理与句法结构紧密耦合,干预会持续降低LaTeX和boxed解格式的质量;(ii)思考模式在受到干扰时会以元认知线索增加、重复且低信息量的延续为特征的补偿性过度生成作出响应;(iii)连贯的CoT行为依赖于专门特征间脆弱的协调,在扰动下会产生不同的失败模式,但输出结构始终受损。
英文摘要
While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit Thinking mode from direct answer generation (NoThinking mode) remain poorly understood. To deconstruct this cognitive process, we apply Top-K Sparse Autoencoders (SAEs) to the intermediate representations of DeepSeek-R1-Distill-Qwen-7B and examine the model's divergent behaviors across math-solving tasks of three distinct difficulty levels. Observationally, we identify a clear distinction in how the model functions under two reasoning modes: Thinking mode relies on sparse and high-intensity feature activations driving verbal deduction independent of problem complexity, whereas NoThinking mode exhibits an adaptive and diffuse pattern prioritizing symbolic manipulation. Causally, suppressing the three most active sparse features by Total Activation Volume reveals three principles: (i) reasoning and syntactic structure are tightly coupled, as interventions consistently degrade \LaTeX{} and boxed-solution formatting; (ii) Thinking responds to disruption with compensatory over-generation marked by increased metacognitive cues and repetitive, low-information continuations; and (iii) coherent CoT behavior depends on a fragile coordination among specialized features, yielding distinct failure modes under perturbation but a consistently impaired output structure.